Etter Solutions LLC

Principal AI Scientist & Consultant

David EtterPh.D.

Vision-language models for event understanding in raw video and multilingual OCR at scale.

David Etter is an applied AI/ML research scientist with over 25 years of experience partnering with industry and academic labs. Since 2006, he has operated Etter Solutions LLC, leading technical efforts across large-scale video retrieval, vision-language model training, optical character recognition across diverse writing systems, and visual anomaly detection. His recent work focuses on high-throughput data pipelines, multimodal retrieval-augmented generation (RAG) over raw video feeds, and robust evaluation benchmarks.

Portrait of David Etter

Train and fine-tune

Large vision-language models, trained with distributed GPUs and fine-tuned for downstream tasks such as OCR and video retrieval.

Evaluate

Benchmarks and test collections built for the task, scored with measures such as character error rate, nDCG, and DET curves.

Deploy

Models prepared for deployment, including export and optimization for fast inference at scale.

Focus areas

Anomaly retrieval in livestream video

In preparation

Thousands of public livestream cameras run continuously around the world. Most of what they capture is static, or shows the typical activity for that place and time of day, and there is far too much of it for a person to watch. This work defines an anomaly as an event that is rare in a camera's own history, and asks a system to rank each camera's clips by how likely they are to be anomalous, with no query and no predefined list of anomaly types.

MultiVENT-Raw Anomaly is a benchmark for this open-set task, built from livestream cameras captured globally between June 2025 and April 2026. Evaluations with both vision-feature and vision-language model approaches show the task remains hard for current models.

livestream cameras
34
video clips
5,859
human-annotated anomalies
76

Event-centric video retrieval

Most video retrieval benchmarks match short visual descriptions against small collections of edited, English-language clips. The MultiVENT collections are built around real-world events in many languages, where the evidence may sit in the picture, the audio, on-screen text, or the metadata.

Multilingual OCR and visual text

Text recognition that works across scripts, and models that read text as pixels instead of tokens.

  • ICDAR 2023 IAPR Best Paper
    A Hybrid Model for Multilingual OCR

    A transformer encoder-decoder that trains the encoder with a CTC objective and the decoder with cross-entropy. The fast encoder can run on its own, or with the full autoregressive decoder when accuracy matters most. Evaluated on the multilingual CAMIO dataset.

  • EMNLP 2021
    Robust Open-Vocabulary Translation from Visual Text Representations

    Replaces a translation model's subword vocabulary with images of rendered text read through sliding windows. On character-permuted German to English it reaches 25.9 BLEU where subword models fall to 1.9.

  • ICDAR 2019
    A Synthetic Recipe for OCR

    One of four ICDAR 2019 papers, alongside work on speech recognition methods for OCR, script identification, and Chinese and Korean character decomposition.

Face recognition

Deep learning models for fast, scalable, and accurate face recognition across a wide range of camera capture settings. The work covers training data and techniques that reduce model bias across gender and ethnicity, metric-learning loss functions, score calibration, and evaluation on the open-set 1:N protocol of the IARPA Janus IJB-C benchmark. More recent models use Vision Transformers and masked autoencoder pretraining.

SCALE workshops

SCALE, the Summer Camp for Applied Language Exploration, is the summer research workshop of the Johns Hopkins Human Language Technology Center of Excellence. Each year it brings researchers and students together for nine to ten weeks to work on one problem in human language technology.

David has taken part in six SCALE workshops since 2012, as a researcher and as a co-lead.

Publications

Selected papers are listed under the focus areas above. The complete list, with citation counts, is on Google Scholar.

Education

  • 2015
    PhD, Computer ScienceGeorge Mason University
  • 2001
    MS, Computer and Information SciencesHood College
  • 1994
    BA, MathematicsShippensburg University

Contact