
A document that takes 20 seconds to generate can still take 20 minutes to use.
Someone has to provide the right information, read the result, check it against the source, correct what is wrong, approve what matters and deal with the cases the system cannot complete. If the output affects a customer, a financial decision or a public service, that checking may require one of the organisation's most experienced people.
The AI step may be fast. The completed workflow may not be.
This distinction matters because many AI business cases measure the most visible part of the process: how quickly a tool drafts, summarises or responds.
They do not measure all the human work required to turn that output into a dependable result.
At Hivemind, we use a simple measure for that: Net Work Removed.
Net Work Removed = baseline human effort − total post-AI human effort
It shifts the question from How fast was the AI? to something more useful:
How much work does the organisation no longer have to do?
Fast output is not finished work

Consider a simple hypothetical.
A team spends 15 minutes preparing a standard piece of correspondence. An AI assistant produces a draft in 10 seconds. It is tempting to record a saving of almost 15 minutes.
But the employee still spends two minutes preparing the input, five minutes checking the draft, three minutes correcting it and another two minutes approving, escalating or recording the outcome.
The completed process now requires 12 minutes of human effort.
The saving is three minutes, not 14 minutes and 50 seconds.
That may still be worthwhile. Across a high-volume process, three minutes per case can create meaningful capacity. But it is a very different business case from the one suggested by generation speed alone.
The same problem appears when a pilot measures isolated tasks rather than the full path from request to completed outcome.
A summary can be created faster while the employee spends longer locating its errors. A response can be drafted faster while approvals accumulate in a new queue. An automated classification can save effort at the start of a process while creating rework for another team later.
The evidence points both ways
There is good evidence that generative AI can improve productivity.
A large field study of 5,172 customer support agents found that access to an AI assistant increased issues resolved per hour by 15 per cent on average. In a controlled study of professional writing tasks, people using generative AI completed their work about 40 per cent faster and produced higher-quality results.
But the benefit is not uniform.
In an experiment involving 758 consultants, AI users completed suitable tasks 25.1 per cent faster and produced higher-quality work. On a task outside the technology's capability frontier, however, they were 19 percentage points less likely to reach the correct answer. The incorrect answers could still look coherent and persuasive.
The lesson was not that AI works or does not work.
It was that the answer depends on the task.
The Australian Government's Microsoft 365 Copilot evaluation found a similar tension. Sixty-nine per cent of matched survey respondents believed Copilot improved the speed of their work and 61 per cent believed it improved quality. Yet verification and editing reduced some of those gains, up to 7 per cent reported that Copilot added time to particular activities, and only 40 per cent of post-use respondents said they had redirected time into higher-value work.
Those figures require care. The trial was non-randomised, participants were relatively experienced and optimistic about generative AI, and the productivity findings relied largely on self-assessment.
That limitation is part of the point.
Perceived time saved is not the same as measured work removed.
What people feel is not always what the workflow shows
The gap between perception and measurement can be surprisingly large.
In a randomised study by METR, experienced open-source developers used early-2025 AI tools on 246 real tasks in repositories they knew well. The developers expected AI to make them faster. Even after completing the work, they believed it had reduced their time by about 20 per cent.
The measured result went the other way: with AI, they took 19 per cent longer.
That study should not be turned into a universal claim that AI slows software development. It examined a narrow group, a particular kind of work and tools that continue to improve. METR's later research also indicates that newer tools may perform better, although its follow-up data could not support a confident causal estimate.
The durable lesson is simpler:
People are not reliable stopwatches for complex work.
A useful evaluation needs operational measurements, not just confidence, satisfaction or an estimate made at the end of the day.
The hidden cost of checking

Review is often described as if it were a small final step.
In practice, it can become a second production process.
A reviewer must understand the request, reconstruct the relevant context, inspect the output, compare it with source material, recognise subtle omissions, decide whether a correction is safe and remain alert even when most outputs are acceptable.
That work consumes time, but it also consumes attention.
The cost rises when the mistakes are plausible rather than obvious. A visibly broken answer is easy to reject. A fluent answer containing one unsupported assumption may take longer to check than writing the answer from scratch.
There is also an expertise paradox.
The people best able to identify a subtle error are often the people whose time the organisation most hoped to protect.
If senior employees become permanent reviewers of routine AI output, the organisation may have moved work rather than removed it.
Human involvement does not automatically solve this problem. A meta-analysis of 106 experiments found that human-AI combinations performed better than people working alone on average, but worse than whichever of the human or AI performed best alone.
The result does not argue against oversight.
It shows that adding a person to a system does not guarantee that the combined system will be better. The review step itself must be designed.
Measure the completed workflow

A better business case begins with the work as it exists today.
Measure the human effort required to complete a representative case before introducing AI. Then measure every form of human effort that remains after the change.
That is what Net Work Removed is intended to capture:
Net Work Removed = baseline human effort − total post-AI human effort
Post-AI human effort should include:
- preparing, locating and entering information
- reading and checking the output
- correcting, rewriting or regenerating it
- approving, escalating or documenting the result
- handling exceptions and incomplete cases
- repairing downstream mistakes
- maintaining instructions, integrations, controls and evaluations
The denominator matters.
A ten-case demonstration selected for convenience is not evidence about a workflow that handles thousands of varied cases. Measure a defined set of eligible cases and include routine, difficult and incomplete examples in proportions that resemble the real work.
Time is also not enough.
Net Work Removed should sit beside measures of quality, completion, cycle time, throughput, cost and risk. A process that removes five minutes while increasing customer corrections or regulatory exposure has not created a useful saving.
This reflects the Australian Government's AI Technical Standard, which warns that a single metric can create false confidence. Its recommended measures include total time and effort, the number of human interventions, performance, safety, reliability, adoption and qualitative outcomes.
Review can still be the right design

The purpose of Net Work Removed is not to eliminate every human check.
In consequential work, review may be essential.
The question is whether it is deliberate, proportionate and valuable.
An AI system that prepares a complex document for an expert may be worthwhile even when every document is reviewed. It may gather evidence, resolve routine structure and remove repetitive drafting while leaving judgement with the person responsible.
The review cost is justified if the complete process becomes faster or better at an acceptable level of risk.
Blanket review is not the only option.
A workflow can direct low-confidence cases, missing information, unusual values or high-consequence actions to a person while allowing lower-risk cases to proceed within defined limits.
It can show the reviewer the source material, identify what changed and record why the case was escalated. It can also stop when it lacks enough information rather than producing a plausible answer.
Good review design reduces unnecessary checking while making necessary checking easier.
Poor review design adds an approval button and assumes responsibility has been solved.
Put the workflow on trial
Before scaling an AI solution, leaders should be able to answer six questions with evidence.
- What is the baseline?
How much human effort does the current workflow require, from request to accepted outcome? - What counts as finished?
Define the quality, completeness, timeliness and risk standard before testing the technology. - What work remains?
Measure preparation, verification, correction, approval, exceptions, rework and ongoing administration. - Who is doing that work?
A saving in junior effort may not be a saving if it creates a larger demand on scarce senior expertise. - What happens across representative cases?
Test the ordinary work as well as the cases most likely to expose a limitation. - What decision will the evidence support?
Decide in advance what would justify proceeding, redesigning the workflow or stopping.
This is why we begin with the business process rather than the AI product.
A prototype should expose where the system helps, where it creates new work and what remains dependent on people.
The useful result is not a demonstration that AI can generate something.
It is a decision about whether the complete workflow has improved.
The better question
AI can save real time.
It can help people work through information, prepare strong first drafts and resolve routine requests more efficiently. The evidence for those benefits is substantial enough that organisations should test them seriously.
But generation speed is not business value, and a confident estimate is not operational evidence.
If every output enters a queue for a person to reconstruct, verify, correct and approve, the organisation needs to count that work too.
The question is not whether a human remains involved.
It is whether the completed workflow uses human effort where judgement, experience and responsibility genuinely matter.
So do not ask only how much time the AI saved.
Ask how much work the organisation no longer has to do. Ask whether the outcome still meets the required standard. Ask what people can now do with the capacity that remains.
Because if AI produces an answer in seconds but leaves people doing most of the work required to trust and use it, the technology may be impressive.
But what, exactly, have you automated?
Sources and further reading
- Australian Government Digital Transformation Agency, Microsoft 365 Copilot evaluation report
- Australian Government, AI Technical Standard Statement 12
- Brynjolfsson, Li and Raymond, Generative AI at Work
- Dell'Acqua and others, Navigating the Jagged Technological Frontier
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
- Vaccaro, Almaatouq and Malone, When combinations of humans and AI are useful
- Noy and Zhang, Experimental evidence on the productivity effects of generative artificial intelligence
- OECD, The effects of generative AI on productivity, innovation and entrepreneurship
- Workday Australia, The AI Tax