Benchmarking the Latest LLMs — July 2024

less than 1 minute read

Published:

Update July 25th, 2024: Benchmarking Large Language Models 🎖

We benchmarked the latest #LLMs: Llama 3.1 (from Meta), Mistral-Large (Mistral AI), GPT-4o-mini (OpenAI).

Findings:

  1. Llama 3.1 405B as good as GPT-4, Claude 3.5 👆
  2. GPT-4o-mini: has drastic performance reduction 👇
  3. Mistral-Large: good alternative but far behind L3.1-405B

Open-source has caught up to commercial models.

The leaderboard 🏁 is available here: huggingface.co/spaces/colab-potsdam/clem-leaderboard

LLM benchmark results

LLM benchmark results