March 2024: Benchmarking the Latest LLMs with clembench

less than 1 minute read

Published:

πŸ“… March 2024: we benchmarked the latest commercial (from OpenAI, Anthropic, Mistral AI) and open LLMs (available on Hugging Face) for their instruction following and self-play abilities without any human annotation, what we call β€œclembench”.

TL;DR: GPT-4 is still the best πŸ”, Claude-3 is closer than before πŸ‘Œ, open models (e.g. OpenChat) are also pushing upwards too πŸ™Œ.

Check out all results and paper details here: huggingface.co/spaces/colab-potsdam/clem-leaderboard

Clembench March 2024 results