Visual Captions: Augmenting Verbal Communication with On-the-fly Visuals

Video conferencing solutions like Zoom, Google Meet, and Microsoft Teams are becoming increasingly popular for facilitating conversations, and recent advancements such as live captioning help people better understand each other. We believe that the addition of visuals based on the context of conversations could further improve comprehension of complex or unfamiliar concepts. To explore the potential of such capabilities, we conducted a formative study through remote interviews (N=10) and crowdsourced a dataset of over 1500 sentence-visual pairs across a wide range of contexts. These insights informed Visual Captions, a real-time system that integrates with a videoconferencing platform to enrich verbal communication. Visual Captions leverages a fine-tuned large language model to proactively suggest relevant visuals in open-vocabulary conversations. We present the findings from a lab study (N=26) and an in-the-wild case study (N=10), demonstrating how Visual Captions can help improve communication through visual augmentation in various scenarios.

UCLA, Los Angeles, California, United States

Google, Mountain View, California, United States

Google Inc., Mountain View, California, United States

Google Research, Mountain View, California, United States

UCLA, Los Angeles, California, United States

Google, San Francisco, California, United States

https://doi.org/10.1145/3544548.3581566

The ACM CHI Conference on Human Factors in Computing Systems (https://chi2023.acm.org/)

Hall G2

6 件の発表

開始日時2023-04-26 23:30:00

終了日時2023-04-27 00:55:00