Presentation
Interpretable Modeling of Driver Attention Shifts with Vision-Language Model
DescriptionDriver attention is central to safe driving, but most current AI systems represent attention only as heatmaps, which show where a driver looked without explaining why attention shifted. This limits their usefulness for understanding driver situation awareness and for designing transparent driver-assistance systems. This project introduces VISTA, a vision-language modeling framework that converts driver attention shifts into structured, human-readable explanations. Instead of only predicting gaze locations, VISTA describes the driving scene, the driver’s current focus, the likely next focus, and the reason for the attention shift.
Using the Berkeley DeepDrive-Attention dataset, we identify important attention transitions by measuring changes between consecutive gaze maps. These selected moments are paired with image frames and gaze overlays, then converted into four-sentence explanations using GPT-4o with human refinement. A LLaVA-NeXT vision-language model is adapted with efficient fine-tuning to learn this explanation format under limited supervision. Results show that the proposed few-shot model substantially improves over zero-shot and one-shot baselines across lexical, semantic, entity-alignment, and human-evaluation metrics. By grounding driver attention in scene entities and rationales, VISTA offers a more interpretable way to model attention shifts and may support safer, more transparent human-AI collaboration in automated driving.
Using the Berkeley DeepDrive-Attention dataset, we identify important attention transitions by measuring changes between consecutive gaze maps. These selected moments are paired with image frames and gaze overlays, then converted into four-sentence explanations using GPT-4o with human refinement. A LLaVA-NeXT vision-language model is adapted with efficient fine-tuning to learn this explanation format under limited supervision. Results show that the proposed few-shot model substantially improves over zero-shot and one-shot baselines across lexical, semantic, entity-alignment, and human-evaluation metrics. By grounding driver attention in scene entities and rationales, VISTA offers a more interpretable way to model attention shifts and may support safer, more transparent human-AI collaboration in automated driving.
Event Type
Lecture
TimeTuesday, October 20th3:20pm - 3:40pm PDT
Location
Similar Presentations

