Etter Solutions LLC

Principal AI Scientist & Consultant

David EtterPh.D.

Vision-language models for event understanding in raw video and multilingual OCR at scale.

David Etter is an applied AI/ML research scientist with over 25 years of experience partnering with industry and academic labs. Since 2006, he has operated Etter Solutions LLC, leading work on large-scale video retrieval, vision-language model training, multilingual OCR, and visual anomaly detection. His recent work focuses on high-throughput data pipelines, multimodal retrieval-augmented generation (RAG) over raw video, and evaluation benchmarks.

Portrait of David Etter

Train and fine-tune

Large vision-language models, trained with distributed GPUs and fine-tuned for downstream tasks such as OCR and video retrieval.

Evaluate

Benchmarks and test collections built for the task, scored with measures such as character error rate, nDCG, and DET curves.

Deploy

Models prepared for deployment, including export and optimization for fast inference at scale.

Focus areas

Event-centric video retrieval

Most video retrieval benchmarks match short visual descriptions against small collections of edited, English-language clips. The MultiVENT collections are built around real-world events in many languages, where the evidence may sit in the picture, the audio, on-screen text, or the metadata.

  • In preparation
    MultiVENT-Raw Anomaly: find the events that are rare in each camera's own visual history

    Whether a crane on a pier is a routine delivery or a one-off depends on that camera's own weeks of footage, so the task is open set: no anomaly classes and no training split. The test collection is 34 anonymized live-stream cameras, 4,666 videos in 5,859 chunks of up to five minutes, with 76 anomalous chunks, roughly 1 in 77. Each camera is one query, every chunk in its archive gets an anomaly score, and the ranking is scored with nDCG@10 averaged over cameras. The work grew out of the 2026 SCALE workshop.

  • arXiv 2026
    MultiVENT-Raw: A Benchmark for Retrieval and Reasoning over Raw Videos

    Nearly 120,000 mostly raw videos (over 5,300 hours) from phones, hand-held cameras, and CCTV, paired with 130 events and 222 event-centric queries. Supports both retrieval and report generation over footage that has no script, captions, or graphics to lean on.

  • CVPR 2025
    MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval

    More than 218,000 news videos and 3,906 queries targeting specific world events. State-of-the-art vision-language models struggle on it.

  • NeurIPS 2023
    MultiVENT: Multilingual Videos of Events and Aligned Natural Text

    About 2,400 event-centric videos grounded in text documents across five languages (Arabic, Chinese, English, Korean, and Russian), mixing news broadcasts with first-hand footage.

Multilingual OCR and visual text

Text recognition that works across scripts, and models that read text as pixels instead of tokens.

  • ICDAR 2023 IAPR Best Paper
    A Hybrid Model for Multilingual OCR

    A transformer encoder-decoder that trains the encoder with a CTC objective and the decoder with cross-entropy. The fast encoder can run on its own, or with the full autoregressive decoder when accuracy matters most. Evaluated on the multilingual CAMIO dataset.

  • EMNLP 2021
    Robust Open-Vocabulary Translation from Visual Text Representations

    Replaces a translation model's subword vocabulary with images of rendered text read through sliding windows. On character-permuted German to English it reaches 25.9 BLEU where subword models fall to 1.9.

  • ICDAR 2019
    A Synthetic Recipe for OCR

    A study of synthetic training data for OCR on a large multilingual set of unconstrained document images, such as internet memes, scanned web pages, and newspapers. It measures how each attribute of the synthetic data affects model accuracy and turns the results into a recipe for generating it.

Face recognition

Deep learning models for fast, scalable, and accurate face recognition across a wide range of camera capture settings, developed between 2017 and 2023. The work covered training data and techniques that reduce model bias across gender and ethnicity, metric-learning loss functions, score calibration, and evaluation on the open-set 1:N protocol of the IARPA Janus IJB-C benchmark. Later models used Vision Transformers and masked autoencoder pretraining.

Recent projects

Recent work in video retrieval, OCR, and aircraft identification.

  • 2026
    DenseIndex: latent-term indexing of dense video embeddings

    Search over dense video embeddings from a plain Lucene index, with no vector database. A sparse autoencoder is trained on the output of a frozen embedder (SigLIP2 or Qwen3-VL), and its active units become the words of each video chunk, so an ordinary inverted index can store and score them. Over 143,000 video chunks a query takes 43 ms on a single CPU thread, and once the chunks are embedded, indexing all of them takes about three minutes; 72–93% of exact dense-retrieval quality is kept on short queries.

  • 2026
    Aircraft Identification: every aircraft in an airfield clip, with its type, operator, tail number, and phase

    A detector draws a box around every aircraft and stops. This task asks what the aircraft is, who operates it, whether it is taking off or landing, and what its tail number says. The MicroAirfield test collection holds 25 live-stream clips from 22 airports in 8 countries, with 186 annotated aircraft across 30 airframe families and 51 operators, each with a bounding-box tube and slots for family, operator, phase, tail number, and description. Scoring matches predicted aircraft one-to-one to annotated ones, so duplicates and false positives cost precision, then grades each slot.

  • 2026
    Contextual OCR: read, translate, describe, and localize every piece of text in a video

    OCR normally returns a transcript and stops. This task asks for the rest: what the text looked like, what it was written on, which pieces belong together, where it was on screen, and what it says in English. The test collection is 12 dense-text videos, 56 minutes in all, with 2,440 annotated text instances in 12 languages across 6 scripts; more than half the text is non-Latin. Scoring matches each prediction to at most one annotated sign and grades the slots separately: character error rate for the transcript, fuzzy match for the translation, exact match for medium and visibility, and facet sets from a closed vocabulary for appearance. A missed sign counts as a full error.

SCALE workshops

SCALE, the Summer Camp for Applied Language Exploration, is the summer research workshop of the Johns Hopkins Human Language Technology Center of Excellence.

David has taken part in six SCALE workshops since 2012, as a researcher and as a co-lead.

Publications

Selected papers are listed under the focus areas above. The complete list, with citation counts, is on Google Scholar.

Education

  • 2015
    Ph.D., Computer ScienceGeorge Mason University
  • 2001
    M.S., Computer and Information SciencesHood College
  • 1994
    B.A., MathematicsShippensburg University

Contact