OpenAI has changed several evaluation benchmarks for its GPT-6 Astra model since first publishing a blog post announcement mid-afternoon on Sept. 3. In some cases, the numbers on the updated versions showed Astra performing better, while numbers for models from OpenAI’s arch rival Anthropic got worse.
The changes occurred amid an unusual rollout of the blog post. OpenAI originally planned for the post to go live at 2 p.m. ET, but it took almost another two hours before it was widely viewable online.
When OpenAI’s X account tweeted out the blog post at 3:32 p.m., the link was not loading properly, returning an error message. At 3:50 p.m., OpenAI CEO Sam Altman posted the link, writing, “We hit a little snag getting the blog post deployed, but it is really great.” Multiple commenters were still unable to see it, and were getting the same error, as did Fortune. When we checked back about an hour later, it was visible and loading properly.
It turns out OpenaAI actually published the blog shortly after 2pm but retracted it for reason the company said it could not disclose, but which it said were unrelated to the benchmark performance figures. (OpenAI first told us it was a bug in the content management system, and then an internet outage.) Upon republishing the blog, it had different evaluation metrics that seemed to favor Astra—and some figures have continued to change even since then.
The revelation of the changes comes amid intense competition in the AI industry, as companies release updates to their large language models at a frenetic pace, each seeking to pull ahead of the other. The focus on metrics also highlight the challenges of measuring the performance of large language models using standardized benchmark tests and concerns that the specs are prone to manipulation and gamesmanship.
“We care deeply about getting evaluations right,” an OpenAI spokesperson told Fortune. “Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run used in reporting. For our launch blog, we made fixes to ensure the numbers represent our best estimate of available model performance, so that users can make meaningful comparisons.”
Discrepancies between the first and final published blogs—and the numbers are still changing
Among the most notable changes was Astra’s reported hallucination rate. In the first internet archive snapshot of the blog post from 2:23, it was 4.2%. It remained that number for several more snapshots, the last being a fifth at 3:11 p.m. ET—about 10 minutes before OpenAI tweeted out the final version.
But the hallucination rate, along with four other metrics, changed in the sixth archival snapshot of the page taken at 5:20 p.m.—after everyone could likely finally see the blog. It was halved down to 2% for Astra. The scores for Astra’s predecessor, GPT-5.6 Sol, also went down from 12.2% to 9.4%. OpenAI has continued to change this metric; as of this writing, the hallucination rates are back up to their original 4.2% and 12.2%.
OpenAI also seems to have given GPT-5.6 Sol a big boost on its internal version of the ExploitBench cybersecurity evaluation, going from 5.5% in the first version to 11.5% in the later versions. OpenAI said it is currently investigating reverting that number back to 5.5% because it says the 11.5% result reflects a reasoning level that is not commercially available for Sol.
Astra is especially good at mathematics, OpenAI says, a quality the company highlights in the opening paragraph of the announcement page. While that metric did not change in the snapshots for Astra—it stays at 97.6% for the FrontierMath Tier 4 (v2) eval—OpenAI did briefly alter the scores for GPT-5.6 Sol and Anthropic’s latest model, Fable 5.1.
The result of these changes made Astra briefly appear significantly better at math than those two models. In the first snapshot (2:23 p.m. on Sept. 3), Anthropic’s Fable 5.1 model’s score is 87.8%. By 5:17 p.m., it’s dropped nearly 10 percentage points to 78%. Today, it’s back up to 83%. Similarly, GPT-5.6 Sol’s scores go from 83%, down to 80.5%, and back up to 83% today.
The changes in metrics began even before OpenAI first published its blog at 2 p.m. An embargoed pre-publication draft the company provided to Fortune and other media organizations listed Astra’s score on the ARC-AGI-3 evaluation as 98.6%. It’s now 99.99% in the live blog.
“We always verify evals before publication so adjustments between draft and final version are normal,” a company spokesperson said at the time. OpenAI also noted that the creator of the benchmark, the Arc Prize Foundation, found that Astra performed at 99.9% in its independent assessment, provided the model was given a particularly powerful harness (a set of tools the model can use to complete tasks). It performed at 63%—still significantly better than any other AI model currently in public release—when given the benchmark’s standard harness. OpenAI said “things like harness, reasoning level and other factors inform evals.”
“Benchmaxxing”—or improving accuracy?
Different research teams at OpenAI oversee different metrics, and are responsible for calculating and reporting them to a central team to publish. OpenAI is open about the fact that the numbers are achieved under the best possible conditions and may be slightly different from the models available in the production ChatGPT product that most users can access. “Evaluation scores are the maximum at any effort,” reads a disclaimer on the blog. The company includes further caveats on each metric in footnotes.
Accuracy is elusive, as multiple numbers can be considered accurate based on the conditions in which the tests occurred. But some AI experts wonder if there’s also “benchmaxxing” involved. This is a known practice in the AI industry—not just at OpenAI—to maximizing scores by re-running evaluations with different conditions.
“This can be done in a very tight timeframe, and it’s better for their marketing,” said Anka Reuel and Mike Hardy, researchers at the Stanford Intelligent Systems Laboratory and Stanford Trustworthy AI Lab. They also pointed out that the GPT-6 Astra system card, which should contain more technical information on how the evaluations were performed, does not always properly explain them. For the internal hallucination benchmark, for example, the system card provides “barely any details about the evaluation,” they said. “It doesn’t even include the number of test items.”
This re-running of the numbers could be why Astra’s coding capabilities also got a marginal boost in the later versions of the blog post, up from 57.7% to 57.9%. Though it’s a negligible difference, OpenAI seemed to care enough about it to swap in the new and improved number.
Not all changes OpenAI made portrayed Astra more favorably. For example, two Anthropic model scores improve in the different versions of the healthcare-focused eval HealthBench Professional. Claude Fable 5.1 goes from 56.6% to 58.1%, and Opus 5 goes from 54.5% to 56.4%. The scores for models made by other AI companies are usually taken from published leaderboards and do not involve OpenAI itself running assessments on rivals’ models.
Evaluation score debates haunt the AI industry
The question of benchmark accuracy has come up multiple times in the past. In 2025, Meta denied reports that it artificially boosted scores for its Llama 4 model by publishing results from an internal version of the model rather than the one it was making publicly-available. Yann LeCun, the former chief AI scientist at Meta, later admitted that the company had “fudged” the benchmark results. Evaluation metrics also change frequently, as new ones get created. For example, ExploitGym, a cybersecurity benchmark that was at the center of the July incident in which OpenAI’s models went rogue and attacked the company Hugging Face, was created in 2026.
Vincent Sunn Chen, an AI engineer at the Snorkel AI, which helps companies building AI models create and evaluate training data, said that it’s not unusual for benchmark scores to shift in the final hours before a model launches. “It’s usually a function of final launch logistics,” he said in an email. “A benchmark score reflects a specific measurement setup: the model checkpoint, configuration (including how much time and compute the model is allowed), harness, eval/grading configuration (e.g., non-determinism in the judge). All of those are typically still shifting in the final days before a launch, so I’m not surprised that there were some updates.” He said he would like to see industry norms developed that companies should report what has changed about the assessment when a company revises benchmark performance numbers so that researchers can interpret the results more clearly.
Benchmark results matter for several reasons. They are the way AI companies measure progress—but also a way to keep score in the race against competing AI companies. Topping the leaderboards for these evaluations can help AI companies win customers, and in some cases help them hire engineers and researchers. But as this example illustrates, interpreting the benchmark scores can be technically complex, presenting a challenge for companies that want to show off the results to the public in a digestible format. These complexities, as well as confusion over changing metrics and accusations that companies have not been intellectually honest in how they’ve presented the results, could make it difficult for customers and investors to figure out exactly which models are best for which tasks. The confusion could muddy the narrative of having the best models in the market that OpenAI would no doubt like to present ahead of a possible 2027 IPO.
The post OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch appeared first on Fortune.




