VideoPrism
Official repository for "VideoPrism: A Foundational Visual Encoder for Video Understanding" (ICML 2024)
About
VideoPrism is a general-purpose video encoder designed to handle a wide spectrum of video understanding tasks, including classification, retrieval, localization, captioning, and question answering. It is pre-trained on a massive and diverse dataset: 1 billion image-text pairs from WebLI, 36 million high-quality video-text pairs, and 582 million video clips with noisy or machine-generated parallel text (subject to data wipeout). The pre-training approach is designed for these hybrid data, to learn both from video-text pairs and the videos themselves. VideoPrism is fairly easy to adapt to new video understanding tasks, and achieves state-of-the-art performance on 31 out of 33 public video understanding benchmarks using a single frozen model.
Open Source Health
- Stars
- 391
- Forks
- 38
- License
- Apache-2.0
- Last commit
- 3 months ago
Resources & Links
Related Categories
Vendor
Google DeepMind
Open-source projects on GitHub: Sonnet, alphafold3 and Lab
Quick Links
Open Source
More by Google DeepMind
Related Products
Deep-Feature-Flow
Deep Feature Flow for Video Recognition
Top category match
MMAction
An open-source toolbox for action understanding based on PyTorch
Top category match
Temporal Segment Networks
Code & Models for Temporal Segment Networks (TSN) in ECCV 2016
Top category match
Cosmos Curator
Cosmos Curator is a powerful video curation system that processes, analyzes, and organizes video content using advanced AI models and distributed computing.
Top category match
Streaming Vlm
StreamingVLM: Real-Time Understanding for Infinite Video Streams
Top category match