I Have a Stream: Making Self-Supervised Learning Work on Continuous Video
NeurIPS 2026
1Faculty of Electrical Engineering and Computing, University of Zagreb
2Fundamental AI Lab, University of Technology Nuremberg
Abstract
Self-supervised learning draws inspiration from infant visual development, yet standard training pipelines bear little resemblance to it: images are independently sampled and globally shuffled across epochs. To reduce this mismatch, we study self-supervised learning from continuous video streams, where frames are consumed in temporal order using strict sliding-window batches, without global reshuffling or multi-epoch replay. To this end, we construct WT++, a 95-hour urban walking-tour video dataset for streaming pretraining.
Combined with a comprehensive evaluation suite, we find that contrastive and distillation-based methods struggle in this setting, while MAE is more robust but still falls short of standard i.i.d. pretraining. We find that high inter-batch similarity, caused by sliding-window consumption across consecutive batches, does not explain this gap. The main challenge is high intra-batch similarity, where frames within each batch are near-duplicates.
To mitigate this, we propose StreamMAE, which preserves the core MAE reconstruction objective while adapting the input pipeline with stream-aware regularization and motion-biased crop selection. StreamMAE outperforms streaming baselines, matches i.i.d. MAE trained on the same video data, remains competitive with ImageNet-pretrained MAE, and scales positively as the pretraining stream grows from 12 to 95 hours.
The streaming setup
How are batches formed from a stream?
- in this batch, new
- in this batch, repeated
- outside the batch
Caveat. This strip holds 64 frames and restarts once it reaches the end. WT++London runs for 12 hours in a single pass. A frame appears in every batch until the window has slid past it, roughly B/s of them, and never appears again. In our runs B = 2048 and s = 1, or s = 2 for WT++95h.
The streaming data
What is the pretraining data?
WT++ is 58 walking-tour videos, 94.5 hours in total, joined into one ordered stream. Each video is a long walk through a city, so the model sees many different objects, lighting conditions and scenes. The default run uses a single 12-hour video, WT++London.
We concatenate city videos into one longer stream
Each stream is the previous one plus more cities.
Ten of the 58 videos
Ten seconds from each. The model is trained on WT++ at 3.75 FPS.
Why does streaming MAE lag behind i.i.d. MAE?
Streaming MAE is worse than i.i.d. MAE trained on the same video. Compared with i.i.d. sampling, a video stream changes two properties of the batches at the same time.
Intra-batch similarity
How similar the frames within one batch are to each other.
0.004 ImageNetvs0.665 video
Inter-batch similarity
How similar two consecutive batches are. Consecutive sliding windows share most of their frames, so this is high for any stream.
0.325 ImageNetvs1.000 video
Both are mean cosine similarities of DINOv2-L features, measured on i.i.d. ImageNet-1K batches and on the WT++12h stream.
The video stream is high on both, so streaming raises the two similarities together. To find out which one causes the drop in performance, we change them one at a time.
Which one matters: inter- or intra-batch similarity?
Both experiments train the same MAE with a ViT-B/16 encoder. Within an experiment, the runs see the same frames and only the order of the frames changes.
- Experiment 1 raises only inter-batch similarity. ImageNet-1K is shuffled once and read with a sliding window. Consecutive batches now overlap, but the images in each batch are still diverse.
- Experiment 2 raises only intra-batch similarity. On WT++12h, a pre-shuffled stream and the chronological stream both have inter-batch similarity close to 1, but only the chronological stream puts near-duplicate frames in the same batch.
Raising inter-batch similarity from 0.325 to 0.989 did not hurt either benchmark, and Cityscapes improved by 1.3 mIoU. Raising intra-batch similarity from 0.171 to 0.665 cost 10.6 mIoU on Cityscapes and 6.1 on ADE20K. Intra-batch similarity is the cause of the performance degradation.
Fine-tuning varies by ±0.7 mIoU on Cityscapes and ±0.3 on ADE20K across seeds, so the collapse in experiment 2 is far outside run-to-run noise. ViT-S and ImageNet-1K accuracy are in the table view.
How does high intra-batch similarity affect optimization?
Under i.i.d. sampling, the cosine similarity between the gradients of consecutive training steps settles near zero. The bigger a streaming method's distance from that value, the worse it transfers.
Distance from i.i.d. MAE in mean cosine similarity between consecutive-batch gradients. The raw value is shown under each name.
Dense transfer gets worse as the distance grows: StreamMAE 0.03, streaming MAE (baseline) 0.15, Orthogonal-MAE 0.24. StreamMAE brings the gradient behaviour of streaming training closer to i.i.d. MAE.
What does StreamMAE change?
StreamMAE keeps the MAE objective unchanged. It changes the training pipeline to reduce the effect of high intra-batch similarity.
Stream-aware regularization
DataDrop
Take a large window and backpropagate through a random quarter of it. The dropped frames are not forwarded at all.
Color jitter
Perturbs colour, brightness and contrast, so reconstruction depends less on appearance that stays the same across neighbouring frames.
Increased drop path
A higher drop path rate varies the effective depth of the encoder across updates. This reduces overfitting to patterns that repeat within a batch.
Stream-aware cropping
Two-stage cropping
First take a fixed-size region of the frame, then apply the standard MAE random resized crop inside it. Choosing where to crop becomes a separate step.
Motion-biased crop selection
Sample K = 4 candidate regions, score each by its mean patch-wise L1 difference to the previous frame, and keep the highest. It is used with probability 0.5 and needs no optical flow.
Ablation study
Components are added one at a time, ViT-S/16 on WT++12h. The baseline here has no DataDrop.
- Partial
- Full StreamMAE
The full pipeline gives the best result on all three benchmarks, so we keep every component. On Cityscapes it gains +3.0 mIoU over plain streaming MAE (baseline), or +2.5 over the DataDrop baseline used elsewhere.
Does it work?
Everything below is pretrained from scratch and evaluated with full fine-tuning.
Can StreamMAE match i.i.d. MAE on the same video?
ViT-S/16 on WT++12h. The dashed line is i.i.d. MAE on that same video.
- StreamMAE
- i.i.d. reference
- Streaming baseline
Yes. It beats every streaming baseline, and it is slightly better than i.i.d. MAE trained on the same video on all five benchmarks.
Does a larger encoder close the gap on its own?
ViT-B/16 on WT++12h.
- StreamMAE
- i.i.d. reference
- Streaming baseline
No. At ViT-B the gap gets bigger, and streaming MAE (baseline) is nearly 10 mIoU below i.i.d. MAE on Cityscapes. StreamMAE reaches 69.0 there, above i.i.d. MAE at 67.8.
Does more pretraining video keep helping?
StreamMAE on streams from 12 to 95 hours. Hover a point to see both encoders.
- ViT-S/16
- ViT-B/16
Yes. Every benchmark is better at 95 hours than at 12 for both encoders, and all except KITTI improve at every step in between.
How does StreamMAE compare with i.i.d. MAE on ImageNet-1K?
Both trained for the same number of iterations, about 650k, matched to the 95-hour stream. StreamMAE sees only walking-tour video, while the reference sees ImageNet-1K.
- StreamMAE on WT++95h
- i.i.d. MAE on ImageNet-1K
With ViT-S, StreamMAE trained on walking-tour video beats i.i.d. MAE trained on ImageNet-1K on Cityscapes (+2.1 mIoU) and ADE20K (+1.0), and is 0.1 points behind on IN-1K. With ViT-B it is behind on all three, most on ADE20K (3.3 mIoU). Checkpoint averaging brings its Cityscapes score to 75.3, just above the ImageNet-1K reference.
Does StreamMAE work only when pretrained on walking tours?
ViT-B/16 on three other continuous streams. Within each stream, all three methods see the same frames for the same number of iterations.
- StreamMAE
- i.i.d. reference
- streaming MAE (baseline)
No. StreamMAE beats streaming MAE (baseline) on every task for kitchen, dashcam and daily-life video. On HD-EPIC and CROWD it also beats i.i.d. MAE on every task.
Poster
BibTeX
@article{martinovic2026streammae,
title = {I Have a Stream: Making Self-Supervised Learning Work on Continuous Video},
author = {Martinovi{\'c}, Ivan and Knobel, Lukas and Asano, Yuki M.},
journal = {arXiv preprint arXiv:2609.40333},
year = {2026}
}