Using Game Play to Investigate Multimodal and Conversational Grounding — accepted at COLING 2025
Published:
Our paper “Using Game Play to Investigate Multimodal and Conversational Grounding in Large Multimodal Models” has been accepted at #COLING2025.
Contribution: Benchmarking of LLMs on multimodal tasks (images + text) in an interactive setting that doesn’t require any annotated test samples.
Key findings:
- Commercial models are way ahead of open-weight models (43 points).
- Even commercial models struggle with fine-grained tasks, e.g. recognizing Pentomino puzzle pieces and telling them apart (see the image below).
- Map navigation is a challenging task where open-weight models mostly get stuck in loops (see the animation below).
Read the full paper here: aclanthology.org/2025.coling-main.381



