Modern multimodal AI systems respond with either text or full images, but many questions are better answered with lightweight visual annotations: an arrow pointing to a component, a highlighted region, or a diagram sketched over an existing image. Despite growing interest in visual communication for AI, no established benchmarks, metrics, or evaluation protocols exist for measuring whether AI-generated visual responses are effective, appropriate, or preferred over text.
This project addresses that gap through two tightly linked research questions. RQ1 (Visual Response Generation) develops models that produce structured graphical primitives (arrows, highlights, diagrams) overlaid on images alongside textual explanations, moving beyond the current paradigm of generating either text or full images. Building on transformer architectures, the model is trained in a multi-task setting to produce communicative annotations that are spatially grounded and semantically meaningful. RQ2 (Evaluation of Visual Dialogue) establishes the first comprehensive evaluation framework for visual responses in AI. I design contrastive benchmarks that pair text-only, visual-only, and hybrid responses to the same queries, and I develop both automated metrics for structural correctness and spatial alignment and human-study protocols for assessing clarity and communicative effectiveness. To support both RQs, I also create datasets linking visual primitives (sketches, gestures) to semantic operations, extending existing image-editing and VQA benchmarks with sketch-based annotations.
The primary contributions are new evaluation benchmarks and protocols for visual dialogue, curated datasets connecting visual primitives to communicative intent, and a visual response generation system capable of producing lightweight communicative graphics. Together, these provide the data and evaluation foundations needed to rigorously assess progress in visual dialogue and enable more natural, efficient, and inclusive human-AI interaction.