05/08/2026
Beyond Chatbots: How We Measure Progress in Artificial Intelligence
When a new AI model is released, headlines often focus on benchmark scores: “Model X outperforms Model Y.” While these comparisons are useful, the benchmarks themselves have evolved significantly over the past two years.
Rather than measuring how much information a model can recall, many modern evaluations aim to assess reasoning, abstraction, and the ability to solve unfamiliar problems.
ARC-AGI-2: Testing General Reasoning
One of the most discussed benchmarks introduced in 2025 is ARC-AGI-2. Unlike traditional language benchmarks, it presents models with abstract visual reasoning tasks that require identifying patterns and applying them to entirely new situations.
The goal is not to test memorized knowledge but to evaluate whether a model can generalize its reasoning to problems it has never encountered before.
What makes ARC-AGI-2 particularly notable is its difficulty. Even today’s strongest reasoning models achieve only modest scores, illustrating that robust general reasoning remains an unsolved challenge in AI research.
To encourage progress, the ARC Prize competition challenges researchers not only to improve accuracy but also to develop solutions that are computationally efficient and cost-effective.
Humanity’s Last Exam
Another ambitious evaluation is Humanity’s Last Exam, a benchmark containing approximately 2,500 expert-level questions spanning more than one hundred academic and professional disciplines.
Rather than rewarding broad factual knowledge, the benchmark was designed to expose the limits of current AI systems by presenting problems that require deep reasoning, domain expertise, and careful analysis. Initial evaluations suggest that even state-of-the-art models continue to struggle with many of these questions.
Interactive AI Demonstrations
Not every AI evaluation is a formal benchmark. Some projects were created to demonstrate machine learning concepts in an intuitive way and remain popular educational resources today.
Examples include:
• Quick, Draw!, where an AI attempts to recognize sketches in real time.
• AI Duet, which generates musical responses as you play a virtual piano.
• Infinite Drum Machine, allowing users to compose rhythms using thousands of machine-learned everyday sounds.
• Talk to Books, an experimental semantic search tool for exploring large collections of books.
• This Person Does Not Exist, which showcases realistic human faces generated entirely by neural networks.
Although many of these projects were developed years ago, they continue to provide accessible demonstrations of modern AI techniques.
From Benchmarks to Practical Evaluation
Researchers and developers are increasingly interested in evaluating AI systems through realistic, multi-step tasks rather than isolated test questions.
Examples include asking different coding models to develop the same software application, comparing how well AI systems plan and execute complex workflows, or assessing creative problem-solving against human participants under controlled conditions.
These evaluations provide a broader picture of a model’s practical capabilities and limitations than benchmark scores alone.
Why It Matters
As AI systems continue to improve, older benchmarks become less informative because many models eventually achieve near-perfect scores.
This has shifted attention toward more demanding evaluations that emphasize reasoning, planning, abstraction, and real-world problem solving. Benchmarks such as ARC-AGI-2 and Humanity’s Last Exam represent this new generation of testing, helping researchers better understand both the strengths and the current limitations of artificial intelligence.
Progress in AI is no longer measured solely by the size of a model or the amount of data it has seen. Increasingly, the central question is whether these systems can reason, adapt, and solve genuinely novel problems.
:::writing
This version reads more like something you’d see from an AI researcher or a technical publication rather than a social media influencer, while still remaining accessible to a general audience.