Grading the Superhuman: Why Measuring AI Is Harder Than Building It

NTVA og Tekna inviterer til åpent møte om hvorfor det er vanskeligere å evaluere kunstig intelligens enn å bygge den, og hvordan vi kan måle om systemene faktisk fungerer når de nærmer seg og overgår menneskelig nivå.
Møtet belyser behovet for realistisk brukersimulering fremfor overmenneskelige språkmodeller, og forklarer hvorfor menneskelig ekspertise forblir helt avgjørende for å skille reell nytte og korrekthet fra overbevisende illusjoner.
Sammendrag av foredraget:
Language models have become remarkably capable in a remarkably short time, and that capability is commoditizing fast: open-weight models now trail the frontier by months, at a fraction of the cost. What has not commoditized is the ability to tell whether any of it actually works. Modern AI systems are not merely measured by benchmarks—they are built out of them. And as these systems approach and surpass human performance, it becomes an open question how to establish that a system is correct at all, and how to tell whether it is still improving.
This talk begins with the foundations—what a language model is, what turns one into an agent, and what current research on self-improvement can and cannot deliver. We argue that the pattern behind AI's very uneven progress is a simple one: systems improve fastest where we can verify them. Coding has a verifier. Being genuinely useful to a particular person does not.
The rest of the talk is about that missing piece. A helpful assistant has to be personalized, which means it needs an internal model of its user. User simulation is therefore not a matter of convenience but an indispensable tool for building —and measuring— genuinely personalized AI. The obvious approach, using a large language model as the simulated user, does not work: language models suffer from a superhuman bias and they lack the cognitive limitations of the users they are meant to represent. A recent study on simulating the reading comprehension skills of Norwegian fifth-graders illustrates what goes wrong, and what it takes instead to build simulators with realistic limitations rather than superhuman ones.
The talk closes on why, even given good simulators and self-improving agents, humans cannot leave the loop. Telling right from merely convincing is measurement work—perhaps the last human expertise the age of superintelligence will need.
The speaker, NTVA-member Krisztian Balog is a professor at the University of Stavanger, where he leads the Information Access & Artificial Intelligence (IAI) research group, and a staff research scientist at Google DeepMind. He also leads the Language and Personalization work package of the Norwegian Research Center for AI Innovation (NorwAI). His research concerns AI systems that help people find, understand, and act on information—currently focusing on how such systems should be evaluated, how they can model the users they serve, and how they can be made transparent. He has published over 200 papers and three books, cited more than 10,000 times (h-index 55). Balog regularly serves as programme chair and senior programme committee member at the leading conferences on information access and AI (SIGIR, WWW, WSDM, EMNLP, ECIR), and is a past and current coordinator of international benchmarking efforts organized by the US National Institute of Standards and Technology (NIST).
Møtet er gratis og åpent for alle interesserte. Det serveres kaffe, te og lapper før og etter presentasjonen.
Annet interessant
Artikler
Generalsekretæren mener at de unges undring kan være en drivkraft for realfagsrekruttering....
Generalsekretæren ønsker et IAEA for KI. (Lederartikkel i NTVAs nyhetsbrev 8. mars 2026.)
Mens de fleste solcellepaneler er plassert på land, representerer flytende solenergi et spennende...
Nyheter
Vi søker en ny generalsekretær til å lede sekretariatet i Trondheim og koordinere vår omfattende...
Akademiet var representert både under selve gravferden i Oslo domkirke, og ved den påfølgende...
Ved H.M. Kong Harald Vs bortgang ønsker Norges Tekniske Vitenskapsakademi (NTVA) å hedre vår høye...