Event-centric video retrieval
Most video retrieval benchmarks match short visual descriptions against small collections of edited, English-language clips. The MultiVENT collections are built around real-world events in many languages, where the evidence may sit in the picture, the audio, on-screen text, or the metadata.
-
In preparation
MultiVENT-Raw Anomaly: find the events that are rare in each camera's own visual history
Whether a crane on a pier is a routine delivery or a one-off depends on that camera's own weeks of footage, so the task is open set: no anomaly classes and no training split. The test collection is 34 anonymized live-stream cameras, 4,666 videos in 5,859 chunks of up to five minutes, with 76 anomalous chunks, roughly 1 in 77. Each camera is one query, every chunk in its archive gets an anomaly score, and the ranking is scored with nDCG@10 averaged over cameras. The work grew out of the 2026 SCALE workshop.
-
arXiv 2026
MultiVENT-Raw: A Benchmark for Retrieval and Reasoning over Raw Videos
Nearly 120,000 mostly raw videos (over 5,300 hours) from phones, hand-held cameras, and CCTV, paired with 130 events and 222 event-centric queries. Supports both retrieval and report generation over footage that has no script, captions, or graphics to lean on.
-
CVPR 2025
MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval
More than 218,000 news videos and 3,906 queries targeting specific world events. State-of-the-art vision-language models struggle on it.
-
NeurIPS 2023
MultiVENT: Multilingual Videos of Events and Aligned Natural Text
About 2,400 event-centric videos grounded in text documents across five languages (Arabic, Chinese, English, Korean, and Russian), mixing news broadcasts with first-hand footage.