March 2024: Benchmarking the Latest LLMs with clembench
Published:
π March 2024: we benchmarked the latest commercial (from OpenAI, Anthropic, Mistral AI) and open LLMs (available on Hugging Face) for their instruction following and self-play abilities without any human annotation, what we call βclembenchβ.
TL;DR: GPT-4 is still the best π, Claude-3 is closer than before π, open models (e.g. OpenChat) are also pushing upwards too π.
Check out all results and paper details here: huggingface.co/spaces/colab-potsdam/clem-leaderboard

