An AI queue can generate more work than your team can safely finish. Before calling that capacity, measure four loads: machine throughput, human verification, exception handling and capability-building. Only the complete ledger shows whether service capacity actually moved.
Agent demonstrations count what the machine produces. Workforce plans must count work completed to the required standard. Those are different numbers.
Faster generation can move the constraint from drafting to verification, exceptions or specialist judgement. It can also consume the time that should have been reserved for coaching, testing and keeping the team's capability current. Until those loads are visible, "capacity released" is an assumption.
Use a four-load capacity ledger instead. It is a team-level method for planning a defined workflow. It does not redesign the checking control, which is a separate job. It does not solve the enterprise pilot-to-scale operating model either. It answers a narrower question: can this team reliably complete more work without concealing a queue or hollowing out capability?
What belongs in the four-load ledger?
Start with one unit of work that everybody can recognise. It might be a de-identified case chronology ready for review, a first-pass control description, a complaint-theme analysis or a draft service response. Define when that unit enters the workflow and when a human is entitled to call it complete.
Then measure four loads over the same period:
- Machine throughput: inputs accepted, outputs produced, outputs abandoned and processing time. This describes the machine lane. It is not a labour saving.
- Human verification: outputs reviewed under the existing control, review minutes, rejected outputs and rework minutes. Use the review standard already approved for the workflow. This ledger measures its demand; it does not weaken it. This is the AI review tax, made explicit.
- Exceptions: cases the system cannot process safely, conflicting-source cases, low-confidence outputs, escalations, reopened work and the human time needed to resolve them. Track the age of the queue as well as its size.
- Capability-building: protected hours for training, supervised practice, calibration, test-case maintenance and learning from exceptions. This is planned work, not leftover time.
Keep the units honest. Output counts belong in the machine lane. Human minutes belong in the other three. Do not add them into a decorative "productivity score". The useful comparison is whether completed, accepted work rose while human load, quality, risk and queue age remained within the limits your team has set.

Jobs and Skills Australia's national study finds that generative AI is more likely to augment human work than replace it, while reshaping tasks and skill needs (Jobs and Skills Australia). That does not tell a bank, insurer or superannuation team how many people a workflow needs. It supports starting with changed tasks rather than a predetermined role reduction.
Safe Work Australia's current AI and digital-technologies guidance makes two of the hidden loads especially visible. It says automation of routine tasks can leave workers with more complex or cognitively demanding tasks, and that more tasks can become focused on reviewing system output. It also identifies insufficient training, unclear roles and work intensification as potential hazards (Safe Work Australia). A capacity plan that omits exception effort and capability time can therefore misread both delivery and work design; the WHS half of that story is the five-signal pace test.
Use this to convert actual operational data into a first ledger. The team leader and workforce analyst must validate every definition and input; the AI must not recommend staffing changes.
Why does fast generation create hidden queues?
Research shows why leaders should resist a simple conversion from AI performance to workforce capacity.
The 2026 peer-reviewed "Cybernetic Teammate" field experiment involved 791 Procter & Gamble professionals working on product-innovation challenges. Individuals using AI matched the performance of two-person teams without AI, and the researchers found that human judgement retained value in evaluating and selecting ideas (peer-reviewed field experiment). That is evidence that AI can change performance on a bounded set of creative tasks. It is not evidence that one AI-assisted employee replaces a two-person operational team, or that the result transfers to regulated casework.
The peer-reviewed "Jagged Technological Frontier" experiment makes the boundary problem clearer. It involved 758 BCG consultants. Across 18 tasks designed to sit inside the tested model's capability frontier, participants with AI completed 12.2 per cent more tasks and worked 25.1 per cent faster on average, with higher assessed quality. On one deliberately selected task outside that frontier, the AI groups were 19 percentage points less likely to reach the correct answer (peer-reviewed field experiment). The study used consulting tasks, subjective rubrics for many inside-frontier outputs and only one outside-frontier task. It establishes uneven performance, not an exception rate for your workflow.
That limitation is the planning point. Your exception rate must be measured locally. An agent may draft straightforward control narratives quickly, then route source conflicts, unusual products and incomplete evidence to people. If those cases take longer than the routine work used to take, the human queue can grow while the machine dashboard looks healthy.
Watch flow in both directions:
- Arrival rate: machine outputs entering human verification, plus new exceptions.
- Clearance rate: outputs accepted and exceptions fully resolved by people.
If arrivals exceed clearance, unfinished work is accumulating. If the oldest item is getting older, the team is not keeping pace even when the total count looks stable. No external study can set the safe gap for you. Use the service obligations, control requirements and risk appetite that govern the actual workflow.
For APRA-regulated entities, the prudential context reinforces the human boundary. APRA's April 2026 industry letter says AI adoption is moving into areas including claims triage and loan application processing, and expects human involvement and accountability for high-risk decisions (APRA). An agent's output volume does not remove those accountabilities, and someone still needs the authority to stop the model when the limits are breached.
When is capacity actually available?
Release a capacity claim only after the workflow passes four gates over a representative observation period.
Machine gate: output volume and failure behaviour are stable enough to plan, including peak periods and known changes. Verification gate: the approved human review can keep pace without deferred checks, falling quality or unplanned overtime. Exception gate: exception volume, resolution time and oldest-case age are stable within the team's limits. Capability gate: protected learning, calibration and control-maintenance time occurred as planned.
Passing those gates does not dictate a headcount decision. It establishes that a measured workflow has changed in a usable way. Leaders still need to consider demand, service levels, resilience, leave, role breadth, concentration of expertise and organisation-specific workforce processes.
Here is a fictional financial-services example using merge-field placeholders.
Workflow: [CLAIMS_OPERATIONS_TEAM] uses an approved agent to draft de-identified case chronologies. Machine load: [OUTPUT_COUNT] drafts produced; [FAILED_OUTPUT_COUNT] failed or abandoned. Verification load: [REVIEWED_COUNT] drafts reviewed at [MEDIAN_REVIEW_MINUTES] minutes; [REWORK_HOURS] hours of rework. Exception load: [EXCEPTION_COUNT] cases, oldest [OLDEST_EXCEPTION_AGE], consuming [EXCEPTION_HOURS] human hours. Capability load: [PLANNED_CAPABILITY_HOURS] planned; [DELIVERED_CAPABILITY_HOURS] delivered. Human decision: [AUTHORISED_LEADER_ROLE] does not claim released capacity because [FAILED_GATE] did not pass. The workflow scope is adjusted and measured again during [NEXT_PERIOD].
The example refuses a false precision. It shows what is known, which gate failed and who decides what happens next.
Use this to stress-test a proposed volume change. The team leader, risk owner and workforce-planning representative must review the assumptions and make any decision.
Before any capacity claim, check that the workflow boundary is fixed, periods match, abandoned outputs are counted, rework is included, exception age is visible, capability time was actually delivered, and a named person has approved the interpretation. If the answer to one item is no, report a measurement gap rather than a saving.
Do this Monday
- Choose one workflow, not an entire function. Write the start point, completion point, unit of work and human decision boundary in one paragraph.
- Set a baseline from recent actual work. Capture demand, completed units, review time, exceptions, queue age and capability hours before comparing the agent-assisted period.
- Instrument the four loads for one operating cycle. Use system events where available and a simple time sample where they are not. Label estimates and replace them with observations as data improves.
- Hold a 20-minute capacity conversion review. Ask which queue moved, which gate failed and what data is missing. Keep quality, risk and work-design signals beside the volume data.
- Delay the capacity claim. Require all four gates to pass over a representative period chosen by the accountable leader. Record the evidence and the limits of the conclusion. Do not convert a demo result or short trial into a staffing number.
Bottom line
Machine output is not spare human capacity. A credible plan measures throughput, verification, exceptions and capability-building across the same workflow and period. Only when all four gates hold can a leader say usable capacity moved, and even then the evidence does not make the workforce decision. The named human leader retains that judgement and records its basis.
This article is general information and education only. It is not legal, compliance, financial or professional advice. Obligations vary by organisation and circumstance. Verify current requirements against the primary sources cited and seek advice specific to your situation.
References
- Fabrizio Dell'Acqua and colleagues, "The Cybernetic Teammate: A Field Experiment on Generative AI and Teamwork", peer-reviewed article published online 12 June 2026. https://pubsonline.informs.org/doi/10.1287/orsc.2025.20702
- Fabrizio Dell'Acqua and colleagues, "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality", peer-reviewed article published online 11 March 2026. https://pubsonline.informs.org/doi/10.1287/orsc.2025.21838
- Jobs and Skills Australia, "Australia's AI Transition: Jobs, Skills and the Future of Work", final release 30 September 2025. https://www.jobsandskills.gov.au/studies/generative-artificial-intelligence-capacity-study
- Safe Work Australia, "Artificial intelligence (AI) and digital technologies - Managing risks", accessed 31 July 2026. https://www.safeworkaustralia.gov.au/safety-topic/hazards/digital-technologies-ai/managing-risks
- Australian Prudential Regulation Authority, "APRA Letter to Industry on Artificial Intelligence (AI)", 30 April 2026. https://www.apra.gov.au/news-and-publications/apra-letter-industry-artificial-intelligence-ai
TheAICommand. Intelligence, At Your Command.


