Explosive Preprint about the AI "Leaderboard Illusion"

TL;DR: leaderboard rankings are unreliable
This picture is worth a t̶h̶o̶u̶s̶a̶n̶d̶ billion words 😱. Based on an preprint on the ArXiv server (The Leaderboard Illusion) by a team led by Shivalika Singh, Marzieh Fadaee, and Sara Hooker from Cohere and colleagues from Princeton University, Stanford University, University of Waterloo, Massachusetts Institute of Technology, Allen Institute and University of Washington (Yiyang Nan, Alex Wang, Daniel D'souza, Sayash Kapoor, Ahmet Üstün, Sanmi Koyejo, Yuntian Deng, Shayne Longpre, Noah Smith, Beyza Ermis), there are reasons to question the performance of some of the most well-known AI models, particularly proprietary and open-weights models.
Their secret? Teaching to the test, literally, and then withdrawing the models that don't perform well, thereby biasing the leaderboards and pumping hype into the "so-close-to-AGI" craze. 🤯
The figure shows the "number of public models vs. maximum arena score per provider, while marker size indicates the total number of battles played. Proprietary model providers tend to achieve higher leaderboard scores, which appear to correlate with both the number of models they release and the number of Arena battles played. Increased exposure to the Arena (through more models and battles) may confer additional advantages, such as better model selection or adaptation to the evaluation distribution. This figure summarizes publicly disclosed results as of April 23rd, 2025."
The worst part of this practice is that it happens without disclosure: private testing, selective score reporting and data access disparities in complete opacity 💢 .
I am curious to see the reactions of the alleged culprits. Does not smell good👃.
The Beta-Lactam Band and the Bogus Upsides
The @PNAS paper by Reese A. K. Richardson (Northwestern University), Spencer S. Hong (Northwestern University), Jennifer A. Byrne (University of Sydney), Luís A. Nunes Amaral (Northwestern University), “The entities enabling scientific fraud at scale are large, resilient, and growing rapidly” is painting a worrisome picture of scientific publishing.
As the authors state, they uncovered “footprints of activities connected to scientific fraud that extend beyond the production of fake papers to brokerage roles in a widespread network of editors and authors who cooperate to achieve the publication of scientific papers that escape traditional peer-review standards.” In other words, the form and scope of corruption is mind blowing. Many researchers, including myself, have received invitations to accept papers without review in return for a monetary or other reward. I always assumed that this approach did not work, but it is apparently working increasing well. I have also been contacted multiple times by fellow scientists who pointed out the existence of perfectly plagiarized versions of some of my papers. There again, I didn’t worry too much. But this has become common. To evade detection, these papers use “tortured phrases” such as “bogus upsides” to mean “false positives” or “Beta-Lactam Band” to mean “Beta-Lactam Group” (have you ever been to one of their concerts?); they “repurpose” and slightly modify images from other (serious) papers.
Paper mills, basically criminal organizations that “that sell mass-produced low quality and fabricated research articles”, are on their way to outpacing actual scientific publications. Perhaps the most disturbing finding to me, from the authors’ sophisticated exploration and network analyses, is that even some of the best journals are not immune to the corruption, as suggested by “Anomalous Patterns in the Editorial Handling of Problematic Publications”.
This is no longer a fringe issue in science. It is threatening the core of the entire scientific endeavor, where progress comes from shared discoveries. With an increasing reliance on AI to explore science literature, we not only have to worry about hallucinations (to be fair, checking that references do exist should be a low bar) but about fraud as well. Some dedicated AI for Science tools are already assigning reputation ratings to citations, which comes with a mixed bag of consequences but seems to be necessary. But in many cases I have seen a retracted paper still being cited by that AI. Keeping track of retractions should be a priority for these chatbots.
The paper, which is Open Access, has many other insights worth exploring -including drastic differences between disciplines. It is a strong wake-up call for scientists as well as established publishing houses.

https://www.pnas.org/doi/abs/10.1073/pnas.2420092122
https://pubpeer.com/publications/0169B1F41428075F52DFBF0E063A20
If AI alone > (AI + Human) > Human alone, what is that telling us?
Recent studies of AI in medicine, particularly in imagery and diagnostic reasoning, have surfaced an alarming trend if we are to believe in AI as a way to augment human abilities: an expert aided by AI is not as good as AI on its own. The figure here is from @eric topol and @pranav rajpurkar’s blog post (erictopol.substack.com/p/when-doctors-with-ai-…), which includes the text of their @new York times OpEd (www.nytimes.com/2025/02/02/opinion/ai-doctors-…). It shows the performance gap in task performance between AI and AI + physician input, which can sometimes be large. Here is Topol and Rajpurkar’s comment on these findings:
“What explains these counterintuitive findings? They could simply reflect that physicians haven't been well grounded in using A.I., or have "automation neglect" (bias against A.I.), or that the studies are relatively small and contrived—attempts at simulating medical practice but a far cry from the complex, messy world of how we diagnose and care for patients. But there may be a more fundamental consideration: we may need to rethink how we divide responsibilities between human physicians and A.I. systems to achieve the goal of synergy (not just additivity, i.e. 1+1 = 5).”
All of these things may contribute to the observed gap, but their last point subsumes a lot of the issues and deserves our attention: the current state of AI-human interaction leads to sub-optimal outcomes. These findings are about specific clinical or medical management tasks and how they generalize is unknown, but diagnostic reasoning, for example, is a commonly required skill across many disciplines: the ability to accurately assess a situation. The 2024 article (jamanetwork.com/journals/jamanetworkopen/fulla…) by @ethan goh, @robert gallo and colleagues provides an excellent starting point: 50 physicians (family medicine, internal medicine, or emergency medicine) were asked to review 6 clinical vignettes in 60 minutes and “randomized to either access an LLM in addition to conventional diagnostic resources or conventional resources only”. Lots of caveats, obviously, e.g., clinical vignettes are not the same as spending time with a patient. But when given this information, the respective median diagnostic reasoning scores for physician with conventional resources, physician with conventional resources and LLM and LLM alone are 74%, 76% and 90%. In other words, the LLM “augmented” the physicians by a small amount (2%). But it could also be said that the physician diminished the LLM’s scores by 14%. For that to happen, the physicians had to dismiss the LLM’s output and favor their own in situations where they differed. The authors of that paper suggest “that further development in human-computer interactions is needed to realize the potential of AI in clinical decision support systems.” Indeed.
The reason I like the diagnostic reasoning case, in addition to its general applicability, is that its psychology has been studied for a long time in the medical space. Numerous cognitive biases have been identified, e.g., availability and self-confidence, as playing an outsized role in diagnostic errors. While self-confidence is something to work on in medical training, availability biases should be addressable by better AI-human interfaces that promote exploration progressively further away from the physician’s comfort/familiarity zone. The progressive aspect is essential for the physician to keep an open mind: anything that is too far from their comfortable center will be perceived as spurious, ridiculous, noise. Think of it as levels in a game: you can only play the next level if you master the current one.
@spencer dorn, @chris cassel, @robert watcher, @jay Parkinson, @thomas wolf, @cassie, @anorld Milstein, @jason maude
https://www.nytimes.com/2025/05/14/technology/ai-jobs-radiologists-mayo-clinic.html
