Benchmarking the Latest LLMs — July 2024
Published:
Update July 25th, 2024: Benchmarking Large Language Models 🎖
We benchmarked the latest #LLMs: Llama 3.1 (from Meta), Mistral-Large (Mistral AI), GPT-4o-mini (OpenAI).
Findings:
- Llama 3.1 405B as good as GPT-4, Claude 3.5 👆
- GPT-4o-mini: has drastic performance reduction 👇
- Mistral-Large: good alternative but far behind L3.1-405B
Open-source has caught up to commercial models.
The leaderboard 🏁 is available here: huggingface.co/spaces/colab-potsdam/clem-leaderboard


