A minimum reporting standard for AI evaluations
A short proposal for the least an evaluation result has to disclose before it can be compared with another result or repeated by anyone else.
Notes
Everything published here, most recent first.
A short proposal for the least an evaluation result has to disclose before it can be compared with another result or repeated by anyone else.
Measuring how language-model agents fail over long tasks. Per-step accuracy barely separates the systems; end-to-end success separates them a lot.
The code I use to run and record my own evaluations, released so the results here can be reproduced.
Which methods for looking inside a model produced findings I would still defend, and which did not survive being checked.
What this practice is for, what it will publish, and the things it will not do.
Everything here is also on the RSS feed.