Clembench April 2025 Update
Published:
We benchmarked the latest released models β GPT-4.1, Llama-4 Maverick and Scout, and Olmo-2-32B β and compared them with existing ones.
Published:
We benchmarked the latest released models β GPT-4.1, Llama-4 Maverick and Scout, and Olmo-2-32B β and compared them with existing ones.
Published:
Weβve put online the latest version of our LLM leaderboard, now based on version 2.0 of clembench β the longest-running game-based evaluation of agentic capabilities of LLMs.
Published:
Our paper on benchmarking multimodal LLMs in an interactive setting that doesnβt require any annotated test samples has been accepted at COLING 2025.
Published:
We benchmarked the most recent Multimodal Language Models using our clembench framework across five dialogue games.
Published:
We benchmarked the latest LLMs β Llama 3.1, Mistral-Large, and GPT-4o-mini. Open-source has caught up to commercial models.
Published:
We benchmarked the latest commercial and open LLMs for their instruction following and self-play abilities without any human annotation β what we call clembench.