All insights
    Abstract streaks of purple and magenta light on black

    AI Visibility

    GPT-6 Astra vs Claude Fable 5.1: reading the launch claims

    Yash Malviya, Co-Founder, White Ocean Media · 6 September 2026

    Two frontier models shipped inside 72 hours. Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on 1 September 2026. OpenAI followed with GPT-6 Astra on 3 September, first as a limited preview for trusted partners, then to paid users the next day.

    We are not a benchmarking lab and this is not a model review. It is a look at how two of the most sophisticated communications teams in technology framed their claims, where independent testing agrees with them and where it does not, because the gap between those two things is the whole subject of announcement credibility, and this is an unusually well-documented example of it.

    What actually shipped

    Claude Fable 5.1GPT-6 Astra
    Announced1 September 20263 September 2026
    AvailabilityGenerally available across AWS, Google Cloud and AzurePreview to trusted partners, then paid ChatGPT tiers, API and AWS
    Restricted siblingMythos 5.1: same model, stronger safeguards, trusted-access onlyPublic version restricts some cybersecurity prompts
    List price$10 / M input, $50 / M output$10 / M input, $50 / M output
    Notable pricing moveCache reads cut 75% to $0.25 / MLong-context surcharge

    Anthropic's framing was cost and reliability: roughly 25% cheaper for typical workloads, up to around 45% for highly agentic work, with agentic coding accuracy reported at 55.8% against Fable 5's 42.0%. OpenAI's framing was capability: the "world's most intelligent and aligned" model, state of the art on computer use, browsing, software engineering and science.

    Both framings are defensible. Both are also selective, in the ordinary way every launch is selective.

    The comparison that wasn't quite like-for-like

    OpenAI's announcement positioned Astra as beating its own GPT-5.6 Sol and Anthropic's Claude Fable 5. Fable 5.1 had shipped two days earlier.

    That is not dishonest, and anyone who has run a launch knows why it happens: benchmark suites are locked weeks ahead, and a competitor shipping 48 hours before you is not something a comms calendar absorbs. But it does mean the headline comparison was against a superseded model, and a reader who does not notice the point release draws a conclusion the data does not support.

    The same care is worth applying to the specific figures. OpenAI reported Astra at 72.6% on OSWorld 2.0's offline set against 70.2% for Claude Opus 5, and 92.7% on ScreenSpot-Pro against 87.3% for Fable 5. Those are real results on real benchmarks, and they are computer-use and interface-grounding tasks, which is where Astra is genuinely strong. They are not a general intelligence claim, though the summary language invites reading them as one.

    Where independent measurement disagrees

    Artificial Analysis, which runs models itself rather than reprinting vendor numbers, puts Astra at 61 on its Intelligence Index against 66 for Fable 5.1 at maximum effort, five points behind, and roughly level with Astra's own predecessor.

    On coding agents the picture reverses: Astra scores 67 on that index, approximately level with Claude Opus 5 and Fable 5.

    So both launch narratives survive contact with independent testing, in their own lanes, and neither survives as a general claim of superiority. "The best model" is not a question with an answer; "the better model for agentic coding at this price point" is.

    The harness story, which is the instructive part

    The clearest example of why methodology deserves as much attention as the number came from ARC Prize, who tested Astra on ARC-AGI-3 twice.

    Run inside ARC Prize's own standard harness, Astra scored 62.7%. Run inside OpenAI's provider adapter (which preserves the model's opaque reasoning state between requests, so it resumes its thinking rather than reconstructing it), the same model scored 99.9%, and cost less to run.

    Both numbers are real. They measure different systems: one measures the model, the other measures the model plus a piece of infrastructure that is not part of what most customers can buy. ARC Prize's own position is that only the 62.7% figure permits a fair comparison between vendors.

    If you take one thing from either launch, take this: the number and the conditions that produced it are a single claim, and quoting one without the other is where credibility gets spent.

    What this means for anyone announcing anything

    Very few companies ship frontier models. Almost every company ships announcements, and the same mechanics apply at any scale.

    • Name your comparison set and its date. "Faster than the leading alternative" invites the reader to supply a competitor you did not test against. Naming the version and the date is both more honest and more defensible.
    • Publish the conditions with the result. A number without its methodology is not a weaker claim, it is a different claim. Journalists who cover your category will ask, and increasingly so will the people using assistants to evaluate you.
    • Expect the correction to travel. Both of these launches were independently re-tested within days. That is now the normal case in any category with an interested technical audience, not a frontier-AI peculiarity.
    • Selective is fine; unfalsifiable is not. Leading with your strongest genuine result is ordinary communications. Framing it so no one can check it is the line worth not crossing.

    This is the same logic behind why a press release that earns coverage looks different from one that gets listed, and behind the questions we suggest asking any agency that claims results it will not let you verify, whether in Dubai or anywhere else.

    Why this specifically matters for AI visibility

    There is now no single assistant to optimise for. Fable 5.1 is generally available across all three major clouds; Astra is rolling out through ChatGPT's paid tiers, the API and AWS. Your buyers are asking different systems, trained and retrieved differently, and getting different answers about your category.

    Two consequences follow.

    Checking one assistant is not a measurement. A single query to a single model on a single day tells you almost nothing, which is why it needs doing repeatably across assistants before it counts as data.

    What gets cited is what can be corroborated. Models increasingly weight sources that agree with each other. A claim that appears only in your own materials is a claim with nothing behind it; the same claim reported independently in credible outlets is what an assistant can safely repeat. That mechanism is unglamorous and it has not changed with this release cycle, and it is the entire reason earned coverage does work that advertising cannot.

    Held to the same standard

    Everything above is sourced to the primary announcements and to independent testers who publish their methodology: Anthropic and OpenAI for what was claimed, Artificial Analysis and ARC Prize for what was independently measured. We have not run these benchmarks ourselves and do not claim to have.

    White Ocean Media works on AI visibility as a PR problem rather than a prompt-engineering one, on the basis that assistants cite what credible independent sources have already established. The track record is public, and it should be read with exactly the scepticism this article recommends applying to a benchmark chart.

    DISCOVERY

    Build the authority that AI can see

    See how this connects to AI Visibility, or tell us about your brand and goals directly.