Run the test, step by step
- Pick one real task you do often. An email you write every week, a summary, a meal plan, a spreadsheet formula. Pick one where the answer comes back as text.
- Choose your two models. The model is the engine inside the app. If you already pay, test your paid model against a free app from another company. If you are on a free plan, test two free apps against each other first, and pay for one month only if the task still comes out wrong.
- Send both the same message, word for word, each in a new chat. Use a temporary chat if the app has one, so saved memory does not help one side. Take out anything private first, swapping real names and account details for made-up ones.
- Open a third new chat, in either app, and paste the prompt below, with the two answers labelled A and B, and do not say which model wrote which. A model that wrote neither answer makes the fairest judge. With only two apps, run the prompt in both and see whether the verdicts agree.
- Tell it which answer was the cheaper one when it asks. It then says whether the cheaper model is enough for that task, and which kind of task to test next.
The prompt: judge two answers to the same task
Replace each bracketed line under ABOUT THE TEST with your answer, paste in both answers, then send.
You are judging two answers to the same task, so I can decide whether a cheaper or free AI model is good enough for it or whether the expensive one earns its price. Your verdict feeds a decision about keeping or cancelling a paid plan, so a polite "both are good" wastes my money, and a confident guess about a fact you cannot check could cost me more. Judge only against my own checks, show your reasons, and never guess which model wrote which answer. ABOUT THE TEST (I fill this in once) - The task I gave both models, word for word: [paste the exact message you sent, e.g. "Write a reply to my landlord asking for a repair date for the heating. Under 150 words, firm but polite, mention my emails of 3 and 10 March." Both models must have had the same words] - The material the task was based on, if any: [paste the document, notes or data both models were given, or write "none"] - What a good answer looks like: [three to five checks you would judge it by, e.g. "under 150 words", "asks for a date in the first two lines", "mentions both earlier emails", "sounds like a person, no legal threats"] - Facts that must be right: [every name, date, number, price or rule the answer depends on, e.g. "emails sent 3 and 10 March", "heating broke 28 February". Write "none" if the task has no facts to get wrong, or "I do not know them" if finding the facts is the task] - The format I asked for: [e.g. "an email, no subject line", "a table with three columns", "only the list, no introduction"] - How often I do this kind of task: [e.g. "most days", "about once a month", "this is a one-off"] - What the expensive option costs me: [e.g. "20 a month and I already pay it", "nothing yet, I am deciding whether to upgrade", "nothing, both are free and I only want the better one"] - What goes wrong if the answer is wrong: [e.g. "mild embarrassment", "I send a client the wrong figure", "I could lose money". Be honest, this decides how strict you are] ANSWER A [paste the first answer here, complete and unedited] ANSWER B [paste the second answer here, complete and unedited] I know which model wrote which. Do not ask yet, and do not guess. BEFORE YOU JUDGE If any line above is blank or still in brackets, list what is missing and wait for me. If the two answers look like replies to different tasks, say so and stop. If either answer looks cut off mid-sentence, say so and wait. If the task asked for an image or a file, tell me this prompt only judges text, and stop. If my checks are too vague to score (for example "make it good"), offer three sharper checks based on the task, ask me to pick, and wait. STEP 1: SCORE EACH CHECK Make a table with one row per check. For each answer, mark the check MET, PARTLY MET or NOT MET, and quote the few words from the answer that made you decide. If a check cannot be judged from the text alone, write [CANNOT JUDGE FROM HERE] and say what you would need. STEP 2: CHECK THE FACTS List every factual claim in each answer: names, dates, numbers, prices, rules, quotes. Check each one against the material and the facts I gave you first. Mark each one MATCHES WHAT I GAVE YOU, CONTRADICTS WHAT I GAVE YOU, or CANNOT CHECK FROM HERE. Never mark a fact as right because it sounds right. If the two answers disagree on a fact, put that at the top of this step, since at least one of them is wrong. If I wrote that I do not know the facts, list the five claims each answer most depends on, mark where A and B disagree, and tell me those are mine to check before I trust either. If you can search the web in this chat, ask me before you do, and give the source for anything you look up. STEP 3: CHECK THE FORMAT AND THE INSTRUCTIONS Say whether each answer did exactly what the task asked: the length, the structure, anything I said to leave out, anything added that I did not ask for (an introduction, a sign-off, an explanation before the answer). Small slips count. List each one. STEP 4: WHAT I WOULD HAVE TO FIX For each answer, list the edits I would need to make before I could use it, and give a rough number of minutes, labelled as an estimate. If an answer is usable as it stands, say so. STEP 5: THE VERDICT Give one verdict, with two sentences of reasons taken from the steps above: - A IS BETTER - B IS BETTER - NO REAL DIFFERENCE FOR THIS TASK (choose this only if both answers have the same marks on every check and need about the same fixing time) Do not reward length, confidence or polish that my checks did not ask for. If either answer has a CONTRADICTS, it cannot win unless the other has one too. If what goes wrong is serious and either answer has facts marked CANNOT CHECK FROM HERE, say that price does not settle this one: I need to check those facts myself whichever model I use. CHECK YOUR JUDGING BEFORE YOU ASK ME ANYTHING 1. Every check has a mark and a quote for both answers. 2. Every factual claim in both answers is on the list in Step 2. 3. The verdict follows from the table. If you swapped the labels A and B, would you still reach the same verdict? If not, say so and score again. 4. Name the one judgement you are least sure about, and what would change it. STEP 6: ASK ME WHICH ONE WAS CHEAPER Once that check is done and shown, ask me which answer came from the cheaper model, and wait. Then tell me plainly: - If the cheaper one won or there was no real difference: the cheaper model is enough for this task. - If the expensive one won: whether the gap is worth paying for, using how often I do the task, what the expensive option costs me, the minutes of fixing from Step 4, and what goes wrong if the answer is wrong. Show that reasoning in three lines or fewer. - If both are free: skip the cost reasoning and say which one to use for this task. - Either way: one task proves one task. Name the kind of task I should test next before I cancel or buy anything, based on what this one did not cover (for example facts, long documents, numbers, tone). RULES - Judge only the text in front of you. Never invent a fact, a source or a flaw. - You may have written one of these answers yourself. Do not try to work out which, and judge both only on my checks. - Never recommend a specific plan, price or product. The decision is mine.
How specific the lines should be, for a weekly client update
What a good answer looks like: under 200 words; late items named in the first three lines; one clear ask; no phrases like "I hope this finds you well".
Facts that must be right: the report is due 26 September; the budget figure is 4,200; the client's name is Dana.
What goes wrong if the answer is wrong: I send a client the wrong date.
For a close call, run it again with the two answers swapped. If the verdict flips, treat it as no real difference.
What the headline numbers mean
In September 2026 a chart went round showing a cheap model almost level with the most expensive ones. Here is what it measured.
The chart: OpenDesign Arena (read 19 September 2026) had 13 models build apps, dashboards and landing pages, scored out of 100, mostly on design quality judged by designers. OpenAI's GPT-6 Astra averaged 82.7, DeepSeek V4.1 Flash 81.2 and Claude Fable 5.1 80.3.
Who runs it: OpenDesign, a company that sells design software. Its own page says the results describe prototype design tasks, "not general model capability".
The part the chart left out: the order changes with the task. On dashboards the cheap model scored 76.1 against Astra's 86.3. On landing pages it scored 79.3 and beat Astra's 76.6.
Before you try DeepSeek
- Your chats are stored in China. DeepSeek's privacy policy (updated 10 February 2026) says: "we directly collect, process and store your Personal Data in People's Republic of China". It also gives you the right to opt out of your data being used to train its models.
The honest bit
- The judge can be wrong too. It can only compare the facts against what you gave it, and it marks the rest "cannot check from here". Those are yours to check, whichever model you end up using.
- A model can favour its own writing, and the answer it reads first. That is why the answers go in unlabelled, and why the swap is worth doing.
- The test judges the answers and nothing else. Free plans also cap how many messages you can send and leave out some features, so check the plan page before you cancel.
- It covers one task, one run, one day. Models change and free plans change what they include, so run it again when a price rises or the answers start to slip.
Try it on this week's task
Pick the task you do most, send it to both models today, and paste the two answers into the prompt. Repeat it on two more tasks before you cancel or buy anything. If you are still choosing between the big names, here is how the four newest models compare.
A few quick questions
Can I use DeepSeek V4.1 Flash for free?
Possibly, but nobody can promise it. DeepSeek's launch post of 10 September 2026 says the model is available through its API, which is the route developers use. It does not say which model the free DeepSeek app runs, so test whatever the app gives you and judge the answers, whatever the model is called.
Do I need to pay to run this test?
No. If you already pay, compare your paid model with a free app. If you are on a free plan, compare two free apps first, and pay for a month only if the task still comes out wrong.
Is the most expensive model always the best?
No. Curtis Pyke at Kingy AI ran four technical tasks twice on three models in September 2026. Claude Fable 5.1 ran up the biggest bill and failed two of eight attempts by adding a sentence before the answer when told not to. The cheapest, DeepSeek V4.1 Flash, failed one. He warns that eight attempts cannot rank anyone.