I made my own leaderboard picture from the Arena agent leaderboard.
I made my own leaderboard picture from the Arena agent leaderboard. A few things to note. 1️⃣ The first, and to me one reason why Arena (with Ion Stoica and Anastasios Angelopoulos) deserves a lot of attention, is that this leaderboard is not based on predefined, potentially gameable, benchmarks: users come to the Arena website, define tasks that matter to them and rank the outputs of models in tournaments. So these tasks are neither abstract nor detached from reality, they represent, to an extent that is real but admittedly hard to assess, actual needs. 2️⃣ The second is that the suspended Claude Fable 5, in its neutral (high) configuration, is by far the best and most reliable. Some 16k sessions took place in the short few days when Fable 5 was available, enough to get a sense that this is not a fluke. 3️⃣ The third is that Anthropic and OpenAI are without a doubt the leaders in frontier models, nothing surprising here. But Z.ai's GLM-5.2 (Max), from the might of its 3 days of existence, is really getting close across 13k sessions. Until now when I heard that US frontier models are one year ahead, I tended to agree but no more. Having played a bit with GLM-5.2, I found it exceptionally good. And with Unsloth AI's incredible shrinking (quantization) of GLM-5.2 to a model that can be run locally on a Mac, something you can do with an open-weights model, something is shifting. In conclusion, I want Fable back, I love Arena and I am in awe of Z.ai.
