Benchmarking the Most Recent Multimodal Language Models
Published:
We benchmarked the most recent Multimodal Language Models using our clembench framework. The benchmark includes 5 dialogue games that evaluate:
- negotiation & reasoning about an image
- reference generation
- map/graph navigation and spatial reasoning (this game has 3 versions)
Findings: commercial models are far ahead in “instruction following” (% of played episodes) and “task solving” (quality score).
- The best commercial model: Claude-3.5 (80 clemscore)
- The best open model: InternVL2-26B (37 clemscore)
The leaderboard is here: huggingface.co/spaces/colab-potsdam/clem-leaderboard
The preprint is here: arxiv.org/abs/2406.14035

