Close

Presentation

Structured Human-AI Teaming for UX Heuristic Evaluation with Human-in-the-Loop Supervision
DescriptionUX heuristic evaluation is a core human factors method for identifying usability problems in interactive systems, but traditional expert-driven approaches are constrained by limited expert availability and time demands. Although multimodal large language models enable automated interface screenshot analysis, fully automated approaches raise concerns about unreliable reasoning and limited transparency. This study proposes a structured human-AI teaming architecture reframing heuristic evaluation as a problem of cognitive labor distribution. In the proposed system, two AI agents—the DR Generator and the Heuristic Evaluator—work with a human supervisor through embedded correction and adjustment loops. The DR Generator produces a structured Design Representation from screenshots, and the Heuristic Evaluator applies established usability principles to identify (generate) potential UX issues. The human supervisor retains supervisory authority by correcting the Design Representation and adjusting the generated issues. The system was evaluated across four Android mobile task scenarios under Baseline, with no human involvement, and Human-in-the-Loop conditions. The Human-in-the-Loop condition achieved 100% correctness, 100% relevance, and 76.25% expert endorsement, compared with 58.92%, 87.50%, and 38.10% in the Baseline condition. Although exploratory and based on a limited sample, the findings suggest that structured human supervision can substantially improve the expert alignment of AI-assisted heuristic evaluation.