Claude Fable 5.1 vs GPT-6 Astra, which one to use

OpenAI launched GPT-6 Astra two days after Anthropic launched Claude Fable 5.1, and every comparison you will read is a chart drawn by one of the two companies. Here is what each one is built for, in plain terms, with both official pages linked and dated. Then a test you run on your own real work, because that is the only scoreboard that matters to you.

Claude & ChatGPT

Two frontier models landed within days of each other and the internet spent the week arguing about benchmark rows. You do not need to follow that. You need to know which tab to open tomorrow morning, so here is the honest version, starting with what each company says about its own model.

The two models, side by side

Pulled from each company's official announcement page, checked again on 9 September 2026. Anything not on those pages is not in this table. Prices are per million tokens; a token is roughly a word.

Claude Fable 5.1 GPT-6 Astra
Announced 1 September 2026 3 September 2026
Where you get it Claude apps and Claude Code, plus AWS, Google Cloud and Azure Rolling out to ChatGPT Plus, Pro, Business and Enterprise over the coming days, plus the API, Azure and AWS Bedrock
Developer price $10 per million tokens in, $50 per million out, with re-read context cut by 75 percent to $0.25 $10 per million tokens in, $50 per million out, with a Fast mode at up to twice the speed for twice the price
What its maker claims The most advanced model for coding and knowledge work The most intelligent and aligned model, and the best at using a computer
What it wins on its rival's own chart The reasoning exam with tools, 65.0 against 57.2, and the Artificial Analysis intelligence index, 65.7 against 61.2 Terminal work 57.9 against 55.8, graduate science 96.0 against 93.7, frontier maths 97.6 against 87.8

That last row is the one worth sitting with. Those Claude wins are printed on OpenAI's own launch page, in OpenAI's own table.

Which one for which job

  1. If the job is words, stay in Claude. Long documents, edits, anything where tone matters. Anthropic's page leads with knowledge work and Fable 5.1 holds the reasoning-with-tools row on both companies' charts, so it is the safer default when the output is something a person will read closely.
  2. If the job is a computer doing the clicking, that is what Astra was built for. Filling forms, updating records, working across tabs and apps. OpenAI put computer use at the front of its announcement and claims the best scores there, so if you want something driven rather than written, that is the pitch.
  3. If the job is hard maths or hard science, check which kind of hard. Astra leads the graduate science and frontier maths rows; Fable 5.1 leads the with-tools reasoning exam. They are good at different shapes of difficult, so send the closed-form problem one way and the messy research question the other.
  4. If you are choosing what to pay for, the app matters more than the model. Both cost the same headline rate for developers, and you will spend far more time inside the interface, the file handling and the memory than you will notice the model underneath. Pick the tool you will actually open.
  5. If you are on a free plan, nothing changes for you today. OpenAI's post lists Plus, Pro, Business and Enterprise for the Astra rollout and says nothing about free accounts, so the honest answer is to wait and read the page again in a week rather than upgrade on launch-day excitement.

The Same-Job Test: how to actually tell them apart

Every chart above is a company grading its own homework. Here is how you get a real answer, and it only works if you make it hard: two models this good will both breeze through an easy task, so an easy task tells you nothing. You separate them on the messy, consequential work, tested the way people who judge these models for a living do it. Give yourself twenty to thirty minutes.

1. Pick your hardest real job, one you have already done

Not a toy prompt. Take the messiest, most consequential thing you actually did recently, ideally one with a few moving parts: an email thread and a spreadsheet, a long document you had to turn into a plan, a pile of numbers you had to make a decision from. Use something you have finished, because then you own the answer key and can mark each model against what good actually looked like. The people who test models seriously make the same point: you only learn a model's real worth on a problem you are genuinely wrestling with, not a puzzle it will always pass (why it takes months to tell if a new model is good).

2. Plant one missing fact, then run the identical job in both

Before you paste it, quietly remove one thing the task needs: a real date, a number, a name. This is the sharpest and cheapest way to tell two strong models apart. Then give both tools the same wrapper, the same task and the same context, so the only thing that differs is the model. Here is what the plant looks like: say you already drafted a client email that quoted a deadline. Delete the date before you paste, and watch. The weaker tool fills in a plausible date, the stronger one writes "deadline not given" and asks you for it. That one line is often the whole answer. Paste this above your task in each.

PromptThe Same-Job Test
I am testing you against another AI model on a real, hard piece of my work. Do the task below properly and in full, as if it were going to be used. Do the actual hard version of it, not a simpler nearby one.

Rules for your answer:
1. Do the whole task. Do not explain what you are about to do first.
2. If anything the task needs is missing, a date, a number, a name, a file, list exactly what is missing in one short line and do NOT invent it. Do the best version you can with what you have, and mark the gap.
3. Follow every instruction in the task, including the small ones.
4. Never invent a fact, a number, a name, or a source. If you cannot verify it from what I gave you, write "not given".
5. At the end, add a section called CHECK: the one part you are least sure about, and one thing I should verify myself before I use this.

Here is the task:
[PASTE YOUR HARDEST REAL TASK. Use something you have already done well, so you know what good looks like, and quietly leave out one fact it needs so you can see whether the model invents it or flags it.]

3. Run it a few times, and clock the speed and the length

Paste the same hard task into a fresh chat two or three times in each tool, and watch whether the strong answer holds its shape or was a fluke. One that is brilliant on the first run and a different shape every run after is not one you can lean on. This is a real effect in how the systems run: at the same settings the same request can land differently depending on how the provider batches it that second (the systems reason why), so one strong answer is never proof. On one run in each, note the two things a chat app will actually show you: roughly how many seconds it took to finish, and how long the answer is (drop it into a word counter). The apps hide the token count, but speed is real when you are the one waiting on it.

This is where a cheap headline price can fool you. The unit that bites is cost per finished job, not cost per token, and a model that is cheaper per word can cost more per job if it writes more or has to be re-run. Two traps hide behind those price rows. The makers do not count a token the same way, so those per-token stickers are not a like-for-like comparison. And a short, confident answer can still be expensive: reasoning models bill their hidden thinking as output tokens, so the visible length tells you little, and the effort or reasoning level is the real cost dial. For the figures a chat app will not show you, Artificial Analysis publishes each model's speed, output per task and cost per task.

Score it on eight things, with the names covered

Swap the gut-feel out of five for eight named checks. If you can, cover which tool produced which answer, label them A and B, and score before you look, so the brand name does not tip the scales. Judge what each answer actually got done, not how confidently it said it, because a surer tone often hides the weaker answer.

One expert catch, because plenty of people now hand the scoring to a third model: a model judging other models is a biased instrument, and the failure modes are named and studied. Judges favour the answer shown first, favour the longer answer, and rate their own family higher (position, verbosity and self-enhancement bias). So if you let a model score them, swap which answer is A and which is B and run it twice, treat any win that came with more words as suspect, and never let a model be the judge of its own family. Your own eye, with the tool names covered, carries less of that bias than you would guess.

  1. Did it do the actual hard job? Not a lazier, easier version of it. Weight this one the most, it is where two strong models pull apart.
  2. The planted missing fact: flagged or invented? Did it say the date was not given, or confidently make one up. The cheapest tell of the two.
  3. Every instruction, including the small ones? The format, the length, the one thing you said not to do.
  4. Did it hold up across the runs? The same shape each time, or a different answer on every run.
  5. Ship-ready, or a redo? Light edits and send, versus start again from scratch.
  6. Honest about what it was unsure of? A real CHECK section naming its weak spot, versus smooth false confidence.
  7. Speed. Which one gave you a usable answer faster on the same task.
  8. Value, not just length. Was the longer, slower answer actually better, or just longer for the same result.

Score each answer yes, partly or no, and add up the yeses, leaning on the first two. If it is a tie, pick on speed, cost and which app you would rather sit in all day. Then keep this task. Run every new model against it as it lands, the way serious testers keep one hard problem of their own (Simon Willison runs the same private test on every new model, precisely because it cannot have leaked into the training data), and you build a real memory of what each model does, worth more than any launch chart.

Where the line is

  • The honest bit: every benchmark in this guide comes from one of the two companies, about itself. OpenAI's table is on OpenAI's launch page and Anthropic's table is on Anthropic's. Neither is independently checked, both companies chose which tests to print, and the numbers were true on the day they were published.
  • A benchmark row is softer than it looks. The same model can post very different scores from nothing more than how the question was formatted, and public tests leak into training data, so a high score can mean the model had seen the exam. Two rows a couple of points apart are noise, which is the real reason a private task of your own beats any chart.
  • Treat them as marketing with maths in it. Trust the Same-Job Test on your own work over any chart, including this one.

Try it on one thing tonight

Take the hardest piece of work you finished this week, paste the wrapper and the task into both tools, and score the two answers. One real task and one honest sitting, and you will stop guessing which tab to open. Model names will change again by Christmas; the test does not.

Next, if the price of all this is the part bothering you, read Which AI subscription is actually worth it.