Correction
How to read a model release without the benchmark theater
The earlier article presented advice on reading model releases as statements from OpenAI Research. A research index did not substantiate those statements. We withdraw those attributions.
Anthropic defines evaluation tasks, trials, and graders and explains why multiple trials matter. These definitions help readers inspect a reported result. They do not validate the earlier statements attributed to OpenAI.
This notice replaces the earlier text. The original URL is retained so readers can see the correction. The reference below was checked during this review and is not presented as the original article's source.