Tag: VideoMAE
-
TimeSformer vs VideoMAE: Efficient Video Understanding
TimeSformer cuts video attention cost by 3x vs full spatiotemporal. Compare it to VideoMAE's masked pretraining with real memory benchmarks.
TimeSformer cuts video attention cost by 3x vs full spatiotemporal. Compare it to VideoMAE's masked pretraining with real memory benchmarks.