Using Game Play to Investigate Multimodal and Conversational Grounding — accepted at COLING 2025

less than 1 minute read

Published:

Our paper “Using Game Play to Investigate Multimodal and Conversational Grounding in Large Multimodal Models” has been accepted at #COLING2025.

Contribution: Benchmarking of LLMs on multimodal tasks (images + text) in an interactive setting that doesn’t require any annotated test samples.

Key findings:

  1. Commercial models are way ahead of open-weight models (43 points).
  2. Even commercial models struggle with fine-grained tasks, e.g. recognizing Pentomino puzzle pieces and telling them apart (see the image below).
  3. Map navigation is a challenging task where open-weight models mostly get stuck in loops (see the animation below).

Read the full paper here: aclanthology.org/2025.coling-main.381

Benchmark overview

Pentomino piece recognition

Map navigation