A framework for training and evaluating proactive egocentric spoken assistants that observe first-person video and audio and provide timely spoken guidance without being explicitly asked.
Code, models, and dataset will be released in November 2026.
Wearable augmented reality (AR) assistants are moving toward continuous real-world interaction, where they perceive the user's activity through first-person video and audio and provide timely spoken guidance without being explicitly asked. We introduce EgoVoice, a framework for training and evaluating proactive egocentric spoken assistants. From HoloAssist video recordings of real human instructors, we construct clean audio streams through source separation, text normalization, and speech resynthesis. We fine-tune an omni-modal LLM and further improve its proactive intervention behavior with direct preference optimization. EgoVoice yields clear improvements in intervention timing, content relevance, and human preference over the zero-shot backbone.
Full-length sessions from the benchmark set, showing the complete task with all instructor interventions.
Sessions are randomly selected from the benchmark set and cropped to approximately 60 seconds (similar to training window length).