<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://sherzod-hakimov.github.io/blog/feed.xml" rel="self" type="application/atom+xml" /><link href="https://sherzod-hakimov.github.io/" rel="alternate" type="text/html" /><updated>2026-08-18T12:45:51+02:00</updated><id>https://sherzod-hakimov.github.io/blog/feed.xml</id><title type="html">Sherzod Hakimov | Blogposts</title><subtitle>NLP and machine learning researcher working on multi-turn, multimodal and multilingual tasks, and on evaluating large language models through dialogue games.</subtitle><author><name>Dr. Sherzod Hakimov</name><email>sherzodhakimov (-at-) gmail.com</email></author><entry><title type="html">Digital Sovereignty in EU — The AI Landscape</title><link href="https://sherzod-hakimov.github.io/blog/2026/digital-sovereignty-eu/" rel="alternate" type="text/html" title="Digital Sovereignty in EU — The AI Landscape" /><published>2026-08-02T00:00:00+02:00</published><updated>2026-08-02T00:00:00+02:00</updated><id>https://sherzod-hakimov.github.io/blog/2026/digital-sovereignty-eu</id><content type="html" xml:base="https://sherzod-hakimov.github.io/blog/2026/digital-sovereignty-eu/"><![CDATA[<blockquote>
  <p>Every person may write to the institutions of the Union in one of the languages of the Treaties and must have an answer in the same language.
— <em>EU Charter of Fundamental Rights, Art. 41(4)</em></p>
</blockquote>

<p>As AI systems are being deployed in different services around the globe, we want to <em>investigate</em> whether current popular LLMs are capable of delivering adequate performance for <strong>certain language capabilities across all 24 languages of the European Union</strong> 🇪🇺.</p>

<h2 id="why-this-evaluation">Why this evaluation?</h2>

<p>Most multilingual benchmarks start from an English multiple-choice test, translate it into other languages and score the answers. This measures knowledge recall through a translation layer, not the ability to use a language. <strong>The real use of LLMs is dialogic</strong>: sustained, goal-directed, with each turn depending on what the model just said and what the user requested.</p>

<p>So we evaluated LLMs as <strong>agents playing dialogue games in self-play</strong> — the model has to follow the rules, track the state, and reach a goal, and it is scored programmatically on whether it did. No gold answers, nothing to leak, and the game mechanics are language-agnostic, so the same benchmark ports to a new language by localising prompt and word-list files rather than authoring a new test.</p>

<p>In our multilingual benchmarking project we tested <strong>9 models · 30 languages · 14 tasks (multi-turn dialogue games) · ~108,000 episodes.</strong> All 24 official EU languages plus Arabic, Chinese, Russian, Serbian, Turkish and Ukrainian. To our knowledge this is the largest multi-turn, interactive evaluation of LLMs across languages published so far — earlier versions of this benchmark covered three languages (mainly English, partially German and Italian). The models were chosen by recency, multilingual capability and good performance overall across other benchmarks.</p>

<p>This gives us 108k episodes of models playing conversational games in different languages, with different success rates. We also look at the amount of web data available for the languages we tested, and available economic data on the regions that they serve. Here is what we found.</p>

<h2 id="what-we-found">What we found</h2>

<p><strong>1. There are two tiers, and the distance between them is wide.</strong> In every single one of the 24 official EU languages, both commercial models score above <em>every</em> open-weight model. Not on average — in all 24. The gap runs from 5.6 points in Greek to 35.8 in Irish, averaging 16 points.</p>

<p><img src="/images/blog/digital-sovereignty-2026-08-cover.png" alt="Best commercial model vs best open-weight model, per EU language" /></p>

<p><strong>2. Parity is achievable, but the open-weight models don’t deliver.</strong> Even though Maltese has roughly four orders of magnitude less public web text than English, GPT-5.4 scores 92.0 in Maltese and 91.8 in English. Irish, Estonian and Latvian look much the same. These languages are not intrinsically harder — they are under-served by open-weight models. It seems like <em>commercial model providers have cracked the code of training data for specific languages</em>.</p>

<p><img src="/images/blog/digital-sovereignty-2026-08-webdata.png" alt="Available web data vs performance" /></p>

<p><em>Available web data vs. performance. The web data values are taken from the HPLT-v3 and FineWeb datasets.</em></p>

<p><strong>3. Open models track the data; closed models do not.</strong> For all seven open-weight models, performance follows both the crawled text available in a language and its economic weight. For the two commercial systems, that correlation is essentially zero. Whatever closes the gap is not in the public crawl.</p>

<p><strong>4. Speakers of smaller languages pay twice.</strong> Tokenisers split Finnish, Maltese and Estonian into roughly twice as many tokens per word as English, and APIs bill per token. Pooled across models, the median non-English language <strong>costs 31% more to run and scores 10% lower</strong>. Equal access to the same model, at the same price, still buys unequal service.</p>

<p><img src="/images/blog/digital-sovereignty-2026-08-cost.png" alt="Performance and cost per language, relative to English" /></p>

<p><em>Performance vs. cost across languages in comparison to English.</em></p>

<p><strong>5. Scale is not the answer.</strong> The 675B-parameter Mistral-Large-3 averages 40.5 across the 30 languages — below the far smaller Gemma-4 and Qwen3.6. Throwing parameters at the problem has not worked; targeted language resources are a different lever.</p>

<p><img src="/images/blog/digital-sovereignty-2026-08-scale.png" alt="Model scale vs EU-24 average performance" /></p>

<p><em>Average performance across all tasks and languages.</em></p>

<p><strong>6. Chinese models are not the answer either.</strong> The open-weight models from other regions understandably don’t value all EU-24 languages equally: the Chinese models we tested are good at Chinese and English, mixed on other languages.</p>

<h2 id="what-needs-to-happen">What needs to happen</h2>

<ul>
  <li><strong>Fund language resources, not just models.</strong> The gap is a data gap, and the market will not close it: the languages that need the most work have the smallest open web data to use as the training source. Public broadcast archives, parliamentary records and digitised national collections are assets Europe already owns.</li>
  <li><strong>Make tokeniser coverage a procurement criterion.</strong> Vocabulary allocation is a design decision that directly sets the price a citizen pays per sentence. It is measurable, and it should be specified.</li>
  <li><strong>Evaluate interactively, and across the whole EU-24.</strong> A model that scores well on translated multiple-choice questions in five large languages tells you very little about what a citizen in Malta or Ireland will experience.</li>
  <li><strong>Be honest about the trade-off.</strong> Today the systems that serve the EU’s smaller languages well are closed and non-European. Sovereignty means either accepting that dependence or investing in resources beyond the open web — but not pretending the choice is not there.</li>
</ul>

<hr />

<p><strong>arXiv preprint of the study:</strong> <a href="https://arxiv.org/abs/2608.01395">https://arxiv.org/abs/2608.01395</a></p>

<p><strong>Leaderboard:</strong> <a href="https://clembench.github.io/leaderboard.html">clembench.github.io/leaderboard.html</a></p>

<p><strong>Source code:</strong> <a href="https://github.com/clembench/multilingual">github.com/clembench/multilingual</a></p>

<p><strong>Current model runs &amp; results:</strong> <a href="https://github.com/clembench/clembench-multilingual-runs">github.com/clembench/clembench-multilingual-runs</a></p>]]></content><author><name>Dr. Sherzod Hakimov</name><email>sherzodhakimov (-at-) gmail.com</email></author><summary type="html"><![CDATA[We evaluated 9 LLMs as agents playing dialogue games across all 24 official EU languages plus six others — 108,000 episodes in total. In every single EU language, both commercial models beat every open-weight model.]]></summary></entry><entry><title type="html">Clembench April 2025 Update</title><link href="https://sherzod-hakimov.github.io/blog/2025/clembench-update/" rel="alternate" type="text/html" title="Clembench April 2025 Update" /><published>2025-04-15T00:00:00+02:00</published><updated>2025-04-15T00:00:00+02:00</updated><id>https://sherzod-hakimov.github.io/blog/2025/clembench</id><content type="html" xml:base="https://sherzod-hakimov.github.io/blog/2025/clembench-update/"><![CDATA[<p>We benchmarked the latest released models and compared them with existing ones:</p>

<ul>
  <li><strong>GPT-4.1</strong> (released on April 14th)</li>
  <li><strong>Llama-4 Maverick and Scout</strong> (released on April 5th)</li>
  <li><strong>Olmo-2-32B</strong> (released on March 20th) — fully open source model (code &amp; data)</li>
</ul>

<p><strong>Findings 📌</strong></p>

<p>🔷 GPT-4.1 matches the performance of GPT-4o (🔵 blue box) while the pricing 💸 dropped 20%. So you’re getting the same ✅ quality but slightly 💸 cheaper.</p>

<p>🔻 Llama-4 models lack the performance of Llama-3.1 and 3.3 (🟥 red boxes). The models got worse both in 📋 instruction following and 🧩 task solving.</p>

<p>🟢📉 Olmo-2-32B is one of the few truly open-source models, but it lags in performance compared to even smaller Llama-8B (🟩 green box).</p>

<p>As usual, all results and source code are openly shared here: <a href="https://github.com/clembench/clembench">github.com/clembench/clembench</a></p>

<p><img src="/images/blog/clembench-2025-04.jpg" alt="Clembench April 2025 results" /></p>]]></content><author><name>Dr. Sherzod Hakimov</name><email>sherzodhakimov (-at-) gmail.com</email></author><summary type="html"><![CDATA[We benchmarked the latest released models — GPT-4.1, Llama-4 Maverick and Scout, and Olmo-2-32B — and compared them with existing ones.]]></summary></entry><entry><title type="html">It is finally here! Clembench Leaderboard v2.0 is live 🎉</title><link href="https://sherzod-hakimov.github.io/blog/2025/clembench-leaderboard/" rel="alternate" type="text/html" title="It is finally here! Clembench Leaderboard v2.0 is live 🎉" /><published>2025-03-03T00:00:00+01:00</published><updated>2025-03-03T00:00:00+01:00</updated><id>https://sherzod-hakimov.github.io/blog/2025/clembench-leaderboard</id><content type="html" xml:base="https://sherzod-hakimov.github.io/blog/2025/clembench-leaderboard/"><![CDATA[<p>We’ve just put online the latest version of our LLM leaderboard, now based on version 2.0 of clembench (“the longest-running game-based evaluation of agentic capabilities of LLMs”).</p>

<p><strong>Highlights ✨</strong></p>

<ul>
  <li>even more games (total: 14): 20 questions, codenames, text adventure, map navigation are added to the old favourites wordle, taboo, reference game, drawing game etc.</li>
  <li>even more instances (total: 817)</li>
  <li>even for the games shared with the previous version (1.6), we generated new instances for 2.0, which means that data contamination is not an issue</li>
  <li>reasoning models added (o3-mini, claude-3.7, deepseek-r1)</li>
</ul>

<p><strong>Reminder 📝</strong></p>

<ul>
  <li>Our “clemscore” is a combination of how well the models follow the game instructions (% played) and how well the concluded game instances were played (quality score).</li>
  <li>We let the same model play both sides of the game (but independently, of course), and we judge the model by the combined performance.</li>
</ul>

<p><strong>Settings for benchmarked models ⚙️</strong></p>

<ul>
  <li>temp=0, max_new_tokens: 300</li>
  <li>Claude-3.7 reasoning budget: 4k</li>
  <li>Deepseek-r1: max_new_tokens limit is removed.</li>
  <li>API backends: OpenAI, Google, Anthropic, OpenRouter (for DeepSeek AI &amp; Mistral AI models)</li>
  <li>Local backend: 2 Nvidia-A100 cluster for remaining open models using Hugging Face library</li>
</ul>

<p><strong>Interesting Findings 💡</strong></p>

<ul>
  <li>Added reasoning time gives models a slight edge, but not terribly much so (o3-mini and claude-3.7 sonnet at around 67, claude-3.5 sonnet and gpt-4o at around 62)</li>
  <li>Deepseek-r1 had problems in returning responses for all queries using OpenRouter</li>
  <li>Reasoning increased latency ⏳: o3-mini vs. gpt-4o =&gt; 7.7 vs. 0.9 seconds, claude-3.7 vs. claude-3.5 =&gt; 5.2 vs. 1.4 seconds; deepseek-r1 vs. deepseek v3 =&gt; 81.7 sec vs. 4.4 seconds</li>
  <li>The closed frontier models are still at the frontier 📈, but not terribly much so (with deepseek-v3 coming in at 53.3 and Llama-3.1-405B at 52.8, to GPT-4o’s 60.5)</li>
  <li>Size doesn’t matter that much anymore (with Llama-3.1-405B only 4 points above Qwen-2.5-72B and Llama-3.1-70B)</li>
  <li>The inventor of all of this, Google (Google DeepMind), still has some catching up to do: Gemini-2.0-flash still trailing OpenAI, Anthropic or even Meta Llama models.</li>
  <li>Europe 🇪🇺: Mistral models unfortunately are not comparable: huge margin between them and other commercial models. Even Llama-70B does much better.</li>
  <li>China 🇨🇳, on the other hand, comes in strongly: Deepseek is mentioned above; Alibaba’s Qwen-Max is a new player that has caught up with the competition.</li>
</ul>

<p>As before, this version will be active for a couple of months, and we will be adding new models as they appear.</p>

<p>On that note: we’re happy to take sponsorship / accept 🤝 compute credit. This surely is getting expensive; cost 💸 : ~500 EUR to benchmark selected models.</p>

<p>Leaderboard: <a href="https://huggingface.co/spaces/colab-potsdam/clem-leaderboard">huggingface.co/spaces/colab-potsdam/clem-leaderboard</a></p>

<p><img src="/images/blog/clembench-leaderboard-2025-03.jpg" alt="Clembench Leaderboard v2.0" /></p>]]></content><author><name>Dr. Sherzod Hakimov</name><email>sherzodhakimov (-at-) gmail.com</email></author><summary type="html"><![CDATA[We've put online the latest version of our LLM leaderboard, now based on version 2.0 of clembench — the longest-running game-based evaluation of agentic capabilities of LLMs.]]></summary></entry><entry><title type="html">Using Game Play to Investigate Multimodal and Conversational Grounding — accepted at COLING 2025</title><link href="https://sherzod-hakimov.github.io/blog/2024/coling-2025/" rel="alternate" type="text/html" title="Using Game Play to Investigate Multimodal and Conversational Grounding — accepted at COLING 2025" /><published>2024-12-16T00:00:00+01:00</published><updated>2024-12-16T00:00:00+01:00</updated><id>https://sherzod-hakimov.github.io/blog/2024/coling-2025</id><content type="html" xml:base="https://sherzod-hakimov.github.io/blog/2024/coling-2025/"><![CDATA[<p>Our paper “Using Game Play to Investigate Multimodal and Conversational Grounding in Large Multimodal Models” has been accepted at <strong>#COLING2025</strong>.</p>

<p><strong>Contribution:</strong> Benchmarking of LLMs on multimodal tasks (images + text) in an interactive setting that doesn’t require any annotated test samples.</p>

<p><strong>Key findings:</strong></p>

<ol>
  <li>Commercial models are way ahead of open-weight models (43 points).</li>
  <li>Even commercial models struggle with fine-grained tasks, e.g. recognizing Pentomino puzzle pieces and telling them apart (see the image below).</li>
  <li>Map navigation is a challenging task where open-weight models mostly get stuck in loops (see the animation below).</li>
</ol>

<p>Read the full paper here: <a href="https://aclanthology.org/2025.coling-main.381/">aclanthology.org/2025.coling-main.381</a></p>

<p><img src="/images/blog/coling2025-1.jpg" alt="Benchmark overview" /></p>

<p><img src="/images/blog/coling2025-2.jpg" alt="Pentomino piece recognition" /></p>

<p><img src="/images/blog/coling2025-map-navigation.gif" alt="Map navigation" /></p>]]></content><author><name>Dr. Sherzod Hakimov</name><email>sherzodhakimov (-at-) gmail.com</email></author><summary type="html"><![CDATA[Our paper on benchmarking multimodal LLMs in an interactive setting that doesn't require any annotated test samples has been accepted at COLING 2025.]]></summary></entry><entry><title type="html">Benchmarking the Most Recent Multimodal Language Models</title><link href="https://sherzod-hakimov.github.io/blog/2024/benchmarking-multimodal-models/" rel="alternate" type="text/html" title="Benchmarking the Most Recent Multimodal Language Models" /><published>2024-09-20T00:00:00+02:00</published><updated>2024-09-20T00:00:00+02:00</updated><id>https://sherzod-hakimov.github.io/blog/2024/benchmarking-multimodal-models</id><content type="html" xml:base="https://sherzod-hakimov.github.io/blog/2024/benchmarking-multimodal-models/"><![CDATA[<p>We benchmarked the most recent Multimodal Language Models using our clembench framework. The benchmark includes 5 dialogue games that evaluate:</p>

<ol>
  <li>negotiation &amp; reasoning about an image</li>
  <li>reference generation</li>
  <li>map/graph navigation and spatial reasoning (this game has 3 versions)</li>
</ol>

<p><strong>Findings:</strong> commercial models are far ahead in “instruction following” (% of played episodes) and “task solving” (quality score).</p>

<ul>
  <li>The best commercial model: <strong>Claude-3.5</strong> (80 clemscore)</li>
  <li>The best open model: <strong>InternVL2-26B</strong> (37 clemscore)</li>
</ul>

<p>The leaderboard is here: <a href="https://huggingface.co/spaces/colab-potsdam/clem-leaderboard">huggingface.co/spaces/colab-potsdam/clem-leaderboard</a></p>

<p>The preprint is here: <a href="https://arxiv.org/abs/2406.14035">arxiv.org/abs/2406.14035</a></p>

<p><img src="/images/blog/multimodal-2024-09.jpg" alt="Multimodal benchmark results" /></p>]]></content><author><name>Dr. Sherzod Hakimov</name><email>sherzodhakimov (-at-) gmail.com</email></author><summary type="html"><![CDATA[We benchmarked the most recent Multimodal Language Models using our clembench framework across five dialogue games.]]></summary></entry><entry><title type="html">Benchmarking the Latest LLMs — July 2024</title><link href="https://sherzod-hakimov.github.io/blog/2024/llms/" rel="alternate" type="text/html" title="Benchmarking the Latest LLMs — July 2024" /><published>2024-07-26T00:00:00+02:00</published><updated>2024-07-26T00:00:00+02:00</updated><id>https://sherzod-hakimov.github.io/blog/2024/llms</id><content type="html" xml:base="https://sherzod-hakimov.github.io/blog/2024/llms/"><![CDATA[<p><strong>Update July 25th, 2024: Benchmarking Large Language Models 🎖</strong></p>

<p>We benchmarked the latest <strong>#LLMs</strong>: Llama 3.1 (from Meta), Mistral-Large (Mistral AI), GPT-4o-mini (OpenAI).</p>

<p><strong>Findings:</strong></p>

<ol>
  <li>Llama 3.1 405B as good as GPT-4, Claude 3.5 👆</li>
  <li>GPT-4o-mini: has drastic performance reduction 👇</li>
  <li>Mistral-Large: good alternative but far behind L3.1-405B</li>
</ol>

<p>Open-source has caught up to commercial models.</p>

<p>The leaderboard 🏁 is available here: <a href="https://huggingface.co/spaces/colab-potsdam/clem-leaderboard">huggingface.co/spaces/colab-potsdam/clem-leaderboard</a></p>

<p><img src="/images/blog/llms-2024-07-1.jpg" alt="LLM benchmark results" /></p>

<p><img src="/images/blog/llms-2024-07-2.jpg" alt="LLM benchmark results" /></p>]]></content><author><name>Dr. Sherzod Hakimov</name><email>sherzodhakimov (-at-) gmail.com</email></author><summary type="html"><![CDATA[We benchmarked the latest LLMs — Llama 3.1, Mistral-Large, and GPT-4o-mini. Open-source has caught up to commercial models.]]></summary></entry><entry><title type="html">March 2024: Benchmarking the Latest LLMs with clembench</title><link href="https://sherzod-hakimov.github.io/blog/2024/benchmarking-latest-models/" rel="alternate" type="text/html" title="March 2024: Benchmarking the Latest LLMs with clembench" /><published>2024-03-14T00:00:00+01:00</published><updated>2024-03-14T00:00:00+01:00</updated><id>https://sherzod-hakimov.github.io/blog/2024/benchmarking-latest-models</id><content type="html" xml:base="https://sherzod-hakimov.github.io/blog/2024/benchmarking-latest-models/"><![CDATA[<p>📅 <strong>March 2024:</strong> we benchmarked the latest commercial (from OpenAI, Anthropic, Mistral AI) and open LLMs (available on Hugging Face) for their instruction following and self-play abilities without any human annotation, what we call “clembench”.</p>

<p><strong>TL;DR:</strong> GPT-4 is still the best 🔝, Claude-3 is closer than before 👌, open models (e.g. OpenChat) are also pushing upwards too 🙌.</p>

<p>Check out all results and paper details here: <a href="https://huggingface.co/spaces/colab-potsdam/clem-leaderboard">huggingface.co/spaces/colab-potsdam/clem-leaderboard</a></p>

<p><img src="/images/blog/clembench-2024-03.jpg" alt="Clembench March 2024 results" /></p>]]></content><author><name>Dr. Sherzod Hakimov</name><email>sherzodhakimov (-at-) gmail.com</email></author><summary type="html"><![CDATA[We benchmarked the latest commercial and open LLMs for their instruction following and self-play abilities without any human annotation — what we call clembench.]]></summary></entry></feed>