French Government Created LLM Leaderboard 'Rigged' for Mistral (comparia.beta.gouv.fr)

🤖 AI Summary
The French government’s AI model leaderboard now publishes a transparent Bradley–Terry (BT) satisfaction ranking — built from thousands of user pairwise votes and reactions on the compar:IA “arena” and developed with the French Center of expertise for digital platform regulation (PEReN). Instead of a naive win-rate, the BT model produces probabilistic scores that account for opponent strength, sparse comparisons and network-wide uncertainty, giving more robust rankings and meaningful confidence intervals even when models haven’t faced every opponent. The platform pairs these satisfaction scores with estimated energy consumption per 1,000 tokens, letting users weigh perceived quality against environmental cost. Energy estimates come from Ecologits’ GenAI Impact methodology (based on ISO 14044 life-cycle principles) and factor in model size, architecture, server location and token output; inference and GPU manufacturing impacts are considered. The leaderboard highlights architecture-driven efficiency differences — for example, a dense Llama 3 405B can consume about ten times more energy than an MoE GLM 4.5 of similar nominal scale because MoE activates only a subset of experts (GLM’s 32B active experts vs full dense activation). Proprietary models are omitted when publishers don’t disclose size/architecture, reinforcing the platform’s push for transparency and encouraging development of both higher-performing and more energy‑responsible models.
Loading comments...
loading comments...