Time saved is how long a task took the old way, minus everything the AI way costs you, from writing the prompt and waiting to reading the answer and fixing it.
Big companies rarely publish a number either: in its markets report of 30 September 2026, the investment firm a16z says nearly 30% of S&P 500 companies report some measurable effect from AI, yet only about 2% report any figure they actually track.
The prompt: test whether AI saves me time
Paste this into ChatGPT, Claude, Gemini or any chatbot and fill in the About me lines. It walks you through the set-up, a mid-week check and the verdicts.
You are my honest time auditor: the one voice in my working week with no reason to flatter the tools I pay for, whose job is to show me which minutes AI gives back and which it quietly takes. I use AI for parts of my work and I believe it saves me time, but I have never checked. Over the next week I will log a few real tasks, and at the end you will tell me, task by task, whether AI is saving me time, whether the way I use it needs changing, or whether I should go back to doing it myself. A task that comes back "drop" is a useful result, because it hands me back time and maybe money. Think in stages, show your working, and ask before you assume. RULES FOR THIS WHOLE CONVERSATION - Never invent a time, a count or a cost. If a number is missing, write "not logged" and say how that changes the result. - Keep my guesses and my measured times apart. Anything I estimated from memory is labelled "estimate" everywhere it appears. - Never round in AI's favour. When a figure is borderline, say it is borderline. - Time spent checking AI's work counts as AI time, and so does time spent fixing a mistake I only found later. - If a line below still shows its bracketed example, treat it as not stated. ABOUT ME (I fill this in once; any line can say "not stated") - What I do: [e.g. "freelance bookkeeper with eight small-business clients", "office manager at a ten-person dental practice", or "I run a small online shop on my own"] - The AI tools I use and what I pay each month: [e.g. "ChatGPT Plus, 20 a month in my currency, plus the free AI built into my email". Write "free" for anything free.] - What an hour of my time is worth, if I want a money answer: [e.g. "45, my client rate", "about 25, my salary divided by my hours", or "skip the money part"] - Tasks I think AI saves me time on, and roughly how often I do each: [e.g. "client follow-up emails, about ten a week; summarising meeting notes, three a week; checking a supplier invoice against the order, twice a week; writing product descriptions, five a week"] - Who is logging: [e.g. "just me", or "me and two colleagues, each with our own log"] - Where I will keep the log: [e.g. "the notes app on my phone", "a spreadsheet", or "a notebook on my desk"] WHICH STAGE WE ARE IN Use the first line that fits. - If I have not pasted a log, run THE SET-UP. - If I paste a log and say the week is not over, run THE MID-WEEK CHECK, even in a second week. - If I paste a log together with last week's card or results, run THE SECOND WEEK. - If I paste a log, run THE WEEK'S REVIEW. If my tracking card from the set-up is missing, first ask me for each task's old-way time and what counts as usable, and wait for my answer before you work anything out. THE SET-UP Step 1: Understand my week. Read what I filled in. Ask me up to five questions, only ones whose answers would change which tasks we pick or how we time them, and offer choices where you can, e.g. "Does a follow-up email start when you open the client's message, or when you start typing?", "Do you still do any of these without AI sometimes: yes or no?". Then stop and wait. If nothing important is missing, say so and carry on. After Step 1, put Steps 2 to 4 to me in one message as a short numbered list I can answer in one go, then stop and wait. Give me the tracking card only once I have answered. Step 2: Pick three to five tasks with me. A good test task happens at least three times this week (two AI runs plus one without), has a clear start and end, and is something I would do anyway. Aim for a mix across writing, summarising or sorting, research or look-ups, and admin or data, and include one task I suspect AI is not helping with. If I listed more than five, recommend which to keep and why. If I listed fewer than three, suggest a few from my work and let me choose. Never add a task without my yes. Step 3: Set the old-way time. Raise the points below only where they apply to my tasks. For each task, ask how long one run takes me without AI. If I have timed it, use that. If I am guessing, mark it "estimate" and suggest I do one run this week without AI and time it, because a time remembered from before I used AI can be wrong in either direction. Never fill in an old-way time yourself. If I would not do the task at all without AI, write "would not do it" as the old-way time. - Agree with me where each task starts and ends (e.g. from opening the client's message to pressing send), and use the same points for the old-way time and every run. - If one run can be much bigger than another, agree a size unit (pages, meeting minutes, invoice lines) and set the old-way time per unit. - If an AI agent or a background task does this while I get on with other work, I log only setting it off and checking the result, and put how long it ran in the note. - If AI on a task works in the background with no prompt (suggested replies, autocomplete), I log 0 for prompting and the whole task time for fixing, and we plan two runs with it switched off, if the app allows, to give the old-way time. - If several of us are logging, take an old-way time from each person, because we do not all work at the same speed. - Agree with me how the without-AI runs get picked: before starting a run, I flip a coin (except tasks I would not do without AI) (heads with AI, tails without) until each task has at least one without-AI run, spread across the week rather than all on one day, so I cannot save an easy one for the stopwatch. Step 4: Agree what counts as usable. For each task, write one plain line with me, e.g. "an email I would send after ten minutes of edits or less", "a summary that gets every decision and owner right". This stops me counting a run as a win when the result was not good enough to use. Step 5: Give me my tracking card. One short block for me to save and paste back at the end, containing: each task with its start and end, its old-way time (measured or estimate), its usable line and how often I expect to do it in a week; the coin-flip rule for the without-AI runs; my tool costs and hourly figure if I gave them; and the log format below. Then tell me in one sentence how to log a run in under a minute, and remind me to log the bad runs too, because a log of only the smooth runs flatters the tool. The log format, one line per run: Date | Task | Size (only if the task varies, e.g. "3 pages") | Minutes prompting and waiting, including every retry and re-word (count waiting only if I could not get on with anything else) | Minutes reading, checking and fixing | Minutes on the rest of the task by hand | Result: as is / edits / redid (put the redo minutes here, never in the by-hand column) / gave up / no AI | Note (optional: what went wrong, what was better than my own version, money it saved me directly) A mistake I find later gets its own row: date, task, 0, the fixing minutes, 0, "later fix for the [date] run". For a no-AI run, put the whole time in the by-hand column. If I forgot the stopwatch, I still add the row with "not timed" and the result, so a redo or a give-up still counts. If more than one person is logging, add a Name column. THE WEEK'S REVIEW Step 1: Check the log first. List anything missing or odd: a row without a number (mark it "not logged" and leave it out of the averages), a task with fewer than two runs, a task with no old-way time. A "not timed" row is left out of the averages but still counts in the as is / edits / redid / gave up tally. A row I mark "from memory" counts as an estimate: show it, but leave it out of the averages if the task has at least two timed rows. If I did a task more often than I logged, use only the logged runs and say this week's total is an undercount. If you can see my calendar, email or chat history, use it only to remind me of runs I may have missed, never to fill in minutes. If the log has a without-AI timing for a task, use it in place of the card's estimate and drop the estimate label. If a gap would change a verdict, ask me about it and wait; otherwise carry on and list the gaps in your answer. Step 2: Work out each task, showing the sums. - A no-AI row only sets the old-way time; if there are several, average them, and leave them out of every other sum and count. - AI time for a run is the prompting-and-waiting minutes, plus the reading-checking-fixing minutes, plus the minutes on the rest of the task by hand, plus any later-fix row for that run. - Saving for a run is the old-way time minus that run's AI time. For a run I redid, add my redo minutes if I logged them, or the old-way time if not; a redone run always counts as a miss, whatever the sum. For a run I gave up on, the saving is minus its AI minutes, because nothing got done. - Time saved per run is the average of those savings; time saved this week is their total. A negative number means AI cost me time, and you say so plainly. - If my runs differ a lot in size, work it per unit of the size we agreed, or split the task into small and large runs, and say which you did. - If the last runs are clearly faster than the first, say so and add the average of the last two, labelled, because practice or a saved prompt can make the week's average lag behind where the task is heading. Step 3: Look at quality. For each task, count as is, edits, redid and gave up. A run I marked "edits" that took more fixing than my usable line allows counts as a miss; show it as "edits, over the line", and never quietly relabel my rows. If my notes say a mistake reached someone else (a client, a customer, my manager), name it in that task's verdict line, and any KEEP for that task must say so in the same line, because a fast run that sends out an error has a cost the minutes do not show. Step 4: Give each task one verdict. Count only AI runs here. A miss is a run I redid, gave up on, or used after more fixing than my usable line allows. Apply these in order and use the first that fits: 1. NOT ENOUGH RUNS: fewer than two timed AI runs (a not-timed run still counts as a miss if it was one). 2. DROP: at least two misses, making up half the runs or more; or the time saved per run is a loss bigger than a fifth of the old-way time. If the old-way time is an estimate, say "drop, on an estimate" and tell me to time one run without AI before I cancel anything. 3. TOO CLOSE TO CALL: the old-way time is an estimate and the time saved per run is less than a fifth of it either way. Tell me to time one run without AI and check again. 4. KEEP: the time saved per run is at least a fifth of the old-way time, and no more than one run in four was a miss. If the old-way time is an estimate, say "keep, on an estimate". 5. CHANGE: everything else, including a small loss or small saving on a measured time, and a good saving with too many misses. Name the single most likely fix, using my notes and where the minutes went (if fixing took longer than prompting, the fix is usually in what I give it): a clearer prompt, an example of a good result pasted in, a smaller piece of the task, a saved prompt instead of typing it fresh each time, or a different tool. Say what to try next week. If a task's old-way time is "would not do it", apply rule 1 first, then skip the rest: these minutes are new time, so write "new task" in the time-saved columns and give the total minutes it took this week. Give it WORTH THE MINUTES if most runs met the usable line and my notes say it was worth having, and NOT WORTH THE MINUTES if not. Leave it out of the money saved and list its minutes under the money line as time AI added. If my notes say the AI version was better even though it was slower, give the time verdict anyway and add "your call: slower but better", because that is a fair reason to keep a task and only I can judge it. If several people logged the same task, give a verdict per person and point out where we differ, because the same tool can suit one person and not another. If one person did the AI runs and another gave the old-way time, say the comparison crosses two people and is weaker for it, and use each person's own old-way time wherever they gave one. Step 5: The money line, only if I gave tool costs and an hourly figure. For each KEEP task, and only those, take time saved per run, times how often my card says I do it in a week, times four for a rough month. If I logged fewer runs than my card expects, say whether I did the task less or forgot to log, and ask if you cannot tell. Times my hourly figure, add any money my notes say a task saved me directly, and set the total against what I pay for the tools each month. Show the sum, and label any sum that uses an estimated old-way time. A task with a negative time saved per run never goes into either money line. Then give a second, labelled line that adds only the CHANGE tasks with a positive time saved per run, so I can see what fixing them is worth. If my hourly figure is a share of a salary, call the result "time worth" that amount, because it only becomes money if I bill those hours or put them into paid work. If I pay for a tool that none of the KEEP tasks use, point it out. If I said skip the money part, skip the sum, but still list any money my notes say a task saved me directly. THE MID-WEEK CHECK No verdicts yet. For each task: runs logged so far, runs still needed to reach two, and whether the without-AI run is done. Point out any row missing a number or put in the wrong column (waiting logged as fixing, a redo with no minutes). If a task has not come up at all, ask whether to swap it for another from my list, and wait. If I say I forgot to log some runs, ask me to add them now with "from memory" in the note. Keep the whole check under ten lines. THE SECOND WEEK Run the week's review on the new log, then put last week and this week side by side for each task: runs, time saved per run, misses, verdict. For a CHANGE task, say whether the fix worked, using last week's "how I will know" line. For a task that was too close to call or had too few runs, pool both weeks' runs and say you pooled them, unless I changed the prompt, the tool or the way I do the task in between; then compare the weeks side by side instead. If I tried a different tool on a task, compare the two tools on the same sums. End with one verdict per task I can act on, and say whether any task needs a third week. HOW TO REPORT BACK Work out and re-check every sum, run count and miss count before you write anything below, and show only the checked figures. Never show a table and correct it further down. 1. GAPS IN MY LOG: one line each, or "none". 2. THE RESULTS: a table with task, AI runs logged, old-way time (marked estimate where it is one), average AI time, time saved per run, time saved this week, as is / edits / redid / gave up, misses, and verdict. 3. WHY EACH VERDICT: one plain line per task, quoting my own numbers or notes. 4. WHAT TO CHANGE NEXT WEEK: for each CHANGE task, the one thing to try and how I will know it worked. 5. THE MONEY LINE: the sums, or one line saying why it is skipped. 6. WHAT TO TEST NEXT: for any task that was too close to call or had too few runs, how to settle it in one more week. 7. MY CARD FOR NEXT WEEK: the tracking card again, with each task's AI runs, time saved per run, misses, verdict, the one thing to try and how I will know it worked, plus this week's log rows, ready to paste back with next week's log. CHECK YOUR WORK BEFORE YOU SHOW ME 1. Every number in your answer comes from my log, my tracking card or my set-up answers, or is a sum of them that you showed. 2. Re-add every sum, recount each task's misses against the definition in Step 4, and confirm the table matches. 3. Every estimate is labelled as one, and no verdict leans on an estimate without saying so. 4. Every run I redid or gave up on counts as a miss and is worked out exactly as Step 2 says, and no no-AI row sits in an AI average or run count. 5. Every "edits, over the line" run is counted as a miss, no row of mine is relabelled, and every "would not do it" task is kept out of the money saved. 6. Each verdict follows the order in Step 4. 7. Name the one verdict you are least sure about, and why.
Run the week, step by step
- Pick three to five tasks with the prompt. Choose jobs you do at least three times a week, so each gets two runs with AI and one without, and agree where each starts and ends.
- Let a coin pick the runs you do the old way. Before a run, flip: heads with AI, tails without, until each task has one timed run done by hand. The software consultant James Shore recommends mixing the two kinds of run through the week and letting chance choose, which also stops you saving an easy one for the stopwatch.
- Time every run, the bad ones most of all. Start your phone’s stopwatch with the task, tap lap when you have the answer you will use, and count every retry and re-word. In a BambooHR survey of 1,608 US desk workers published on 1 September 2026, workers said about 42% of their AI time goes on fixing errors and re-prompting.
- Log each run straight away, even when you forgot the stopwatch. A few numbers and a word for how it went take under a minute. A run you forgot to time still goes in as “not timed” with its result, so a redo still counts.
- Mark the runs you redo or give up on, and log later fixes. A redone run costs the AI time plus the minutes you spent redoing it, and a mistake you find on Thursday gets its own row against Monday’s run, so the log counts it the way your week did.
- Paste the log mid-week, then again at the end. Mid-week, the prompt checks for gaps without giving verdicts. At the end you get a results table, a verdict per task, one fix to try for any task marked change, and a card to carry into a second week.
One task’s week, worked through (an example, not real data)
- The task: client follow-up emails, 12 minutes each the old way, timed once on Monday. Usable means an email you would send after ten minutes of fixing or less.
- The log: five runs. Four took 3 + 4, 2 + 9, 3 + 3 and 2 + 5 minutes of prompting plus fixing. On the fifth it kept offering a discount that does not exist, so after 3 + 4 minutes the email was written by hand in 12.
- The sum: the first four saved 5, 1, 6 and 5 minutes. The fifth cost 3 + 4 + 12 = 19, a loss of 7. That makes 10 minutes saved this week, or 2 a run.
- The verdict: change. A keep needs a saving of at least a fifth of the old time, 2.4 minutes here, and the notes point to one fix: paste in the real offer terms so it stops making up a discount.
- What one bad run did: without the fifth run, the saving averages over 4 minutes a run, a keep. One invented discount turned it into a change.
If your work tool already shows “hours saved”
- Microsoft 365 Copilot: its dashboard (read 5 October 2026) credits about 6 minutes for each search, summary or drafting action, plus the time of any meeting it summarises or recaps, and values each hour at $72 by default, so the figure cannot see the minutes you spent fixing.
- ChatGPT Enterprise and Edu: its workspace analytics (read 5 October 2026) show the share of people who say, in a survey, that ChatGPT saves them time.
- Your log: counts the minutes that actually happened, retries and fixes included, so use it to check the dashboard’s figure.
The honest bit
- One week is a small sample. A task you did twice can swing either way, and a tool you only just started is slower while you learn it, so run a second week before you drop a close call.
- Time is one reason to keep a task. You might keep a slower one because the result is better or you hate doing the job, and the prompt flags “slower but better” for you to decide.
- The prompt was run once in ChatGPT Free on 5 October 2026, and two of its sums were wrong. The prompt now tells it to check every sum before it shows you anything, but that version has not been re-run yet, so check its totals against your own log before you drop a tool or cancel a plan.
Start with the next task you do today
Paste the prompt, pick your tasks in about five minutes, flip the coin and time the very next one.
When the verdicts come in, stop using AI for anything marked drop (if it says “on an estimate”, time one run by hand first), save the prompt that worked for each keep, and if the money line shows a plan costing more than it saves, compare what each AI subscription gives you before it renews.
A few quick questions
Does this work for a small team?
Yes. Each person keeps their own log with a Name column and gives their own old-way time, because people work at different speeds. Paste the logs in together and the prompt gives a verdict per person. Use it to judge the way you use AI, never to rank people.
What if a task saves me money rather than time?
Write it in the note column, for example “did not pay a translator this time, saved 40”. The prompt lists it in the money line at the end, even if you skipped the hourly part.