The Estonian government has released a benchmark to determine 'which LLM is best at countering Russian propaganda?'

The Estonian Language Institute has released its 'Propaganda Resistance' benchmark, which measures the resilience of large-scale language models to Russian propaganda. The results showed that Anthropic's Claude Opus 4.7 came out on top overall, with models from NVIDIA and Alibaba also ranking highly.
Propagandakindlus - Keelemudelite mõõdupuu
These LLMs are the best at resisting Russian propaganda - Ars Technica
https://arstechnica.com/ai/2026/06/these-llms-are-the-best-at-resisting-russian-propaganda/
This benchmark measures how well the model can respond to themes used by the Russian Federation in its strategic narratives without being influenced by propaganda. The evaluation consisted of 75 questions in three languages, covering 14 different Russian propaganda narratives.
Questions were categorized into neutral, biased (containing false assumptions), and malicious (attempting to elicit explicit misinformation). Responses were scored from 1 to 5 points, with 5 points indicating a balanced and insightful response, and 1 point indicating a response that merely amplifies propaganda, such as repeating Moscow's arguments.
The results for the top 10 are as follows. Many of Anthropic's Claude series models ranked highly, with six of the top 10 models being from the company's Sonnet and Opus series. The top-ranked Claude Opus 4.7 received the highest rating in 77% of the questions, with an average score of 94.9 out of 100. Openweight models also showed strong results, with NVIDIA's Nemotron 3 Super 120B and Alibaba's Qwen 3.6 Plus ranking highly and performing at a level close to Anthropic's top-of-the-line models.

Among OpenAI's models, GPT-5.4 performed best, receiving the top rating in 54% of the questions and achieving an average score of 88.9. On the other hand, GPT-3.5 Turbo ranked last in the table, highlighting the significant difference compared to older models.
The rankings from 11th to 20th are as follows.

Google's models were found to be vulnerable to malicious prompts and questions in Russian. The Gemini 2.5 Pro scored 66.1 on malicious questions and 75.5 on Russian, while the Gemini 3.5 Flash also scored lower in Russian than in English.

The evaluation used a judgment model that was tuned to closely resemble human experts. The judgment model's evaluation matched the human expert's evaluation by 88% to 100% within 1 point, and the final score was calculated using a geometric mean to ensure that some strengths did not excessively compensate for other weaknesses.
The Estonian Language Institute explained that this benchmark measures the capabilities of the underlying model itself, without using external search, memory, or tools, rather than the overall chatbot user experience.
Related Posts:
in AI, Posted by log1i_yk







