AI Made Drafting Cheap. Build a Rejection Rubric., practitioner guidance from TheAICommand
← Leadership
Leading with AI

AI Made Drafting Cheap. Build a Rejection Rubric.

When a team can generate ten fluent alternatives in minutes, the leadership bottleneck moves from producing options to rejecting them. Lock relevance, evidence, audience and consequence before generation, apply hard reject triggers before any scoring, and record why a survivor deserves further human attention rather than approval.

Leading with AI. Written for Australian managers and people leaders. General information only. The judgement stays yours.

Quick answer

Rejection. When AI can generate ten fluent alternatives in minutes, generation stops being scarce and the bottleneck moves to culling. Lock relevance, evidence, audience and consequence tests before the options arrive, apply hard reject triggers before any scoring, and record why a survivor earned further human attention rather than approval.

When AI can produce ten plausible alternatives before the meeting starts, generation is no longer the scarce skill. Set a reject-first rubric for relevance, evidence, audience and consequence, then record why any surviving option deserves human attention, source checking and further work.

AI made another draft cheap. It did not make choosing one safe.

When your team can generate ten fluent alternatives in minutes, the leadership bottleneck moves to rejection. Which options answer the real question? Which rest on verified evidence? Which work for the intended audience? Which create a consequence the drafter has not confronted?

Set those tests before anyone sees the options. Apply hard reject triggers first, then score only what survives. A person applies the rubric, records the reason and decides whether any option deserves further work. Survival is not approval, and a high score never replaces source checking or an authorised decision.

Why does cheap drafting change the leadership job?

The evidence that AI can compress drafting time is real but bounded. Noy and Zhang ran a preregistered online experiment with 453 college-educated professionals completing occupation-specific, incentivised mid-level writing tasks. Half were randomly given access to ChatGPT. Average completion time fell by 40 per cent and output quality, as assessed in the study, rose by 18 per cent (Science).

That experiment used writing tasks and a 2023 version of ChatGPT. It did not test Australian financial-services decisions, regulatory interpretations or customer outcomes. Its useful leadership signal is narrower: when first-pass production becomes faster, the relative value of framing, selecting and checking increases.

A newer peer-reviewed field experiment makes the selection point explicit. The Cybernetic Teammate involved 791 Procter & Gamble professionals working on real product-innovation challenges. Participants were randomly assigned to work individually or in pairs, with or without AI. Decomposing the innovation process, the authors report that AI primarily enhanced the quality of generated ideas, shifting the distribution of creative output upwards, while human judgement retained value in evaluative selection (INFORMS journal article).

The decomposition is sharper than it first looks. Participants without AI picked their single best idea from five roughly 50 per cent of the time, against roughly 37 per cent for the AI-enabled conditions, although the authors note that the lift AI gave to idea quality more than compensated for that selection gap in the final outcomes. That is one consumer-products setting, not proof that every AI-generated option pool improves.

The common failure is to generate first and invent the criteria later. Once a polished option attracts support, evidence becomes optional, audience fit becomes taste and consequences arrive as late objections.

A rejection rubric reverses that sequence. It does not ask which option is most impressive. It asks which options have already disqualified themselves.

This is not workflow measurement, completed-draft review, red teaming or idea-pool design. It is a narrow gate between generation and serious human attention, reducing the alternatives that receive expensive review without weakening the review survivors still require.

Two boundaries are worth naming, because the same research gets used for different arguments. Applying the same test to the enthusiast and the refuser is a question about people and how you assess them. This is a question about artefacts: the rubric never scores the drafter. And where AI narrows the range of ideas your team generates is a problem with the input pool, a rejection rubric operates after generation and cannot widen a pool that arrived narrow.

What belongs in the REAC rejection rubric?

A style-led shortlist beside a reject-first record with four options culled
Rejection before scoring. Style attracts support; a hard reject ends an option before it can.

REAC stands for Relevance, Evidence, Audience and Consequence. Define each criterion for the specific task before generation. Add a hard reject trigger wherever failure cannot be offset by strength elsewhere.

After the hard-reject pass, use a simple score for the remaining options: zero means the criterion is not met, one means it is partly met with a named gap, and two means it is met on the supplied evidence. Set the survival threshold before scoring. Never average away a hard reject.

Relevance asks whether the option solves the assigned problem. Name the exact question, required outcome, scope, constraints and decision stage. Hard reject an option that answers a different question, ignores a mandatory constraint, introduces an unrequested workforce or customer outcome, or claims authority the exercise does not have.

Evidence asks what carries each material statement. Require the approved source, date or version, and separate fact from assumption. Hard reject invented citations, unsupported claims, a superseded source, a missing controlling record or an inference presented as an established fact. A source-shaped sentence is not evidence, which is why an evidence grade on the update itself does the same work one level up.

Audience asks whether the intended recipient can use the option correctly. Define the reader's knowledge, purpose, required action, accessibility needs and likely misunderstanding. Hard reject language that conceals uncertainty, uses internal shorthand with customers, buries a required action or implies that AI made a judgement reserved for a person.

Consequence asks what happens if the option is used or believed. Identify who may act, what could be irreversible, which control or obligation is engaged, who owns escalation and what downside remains. Hard reject an unauthorised promise, a recommendation beyond the drafter's remit, an unacknowledged material risk or a route that removes required human review.

Set the four out before anyone generates anything:

CriterionThe question it answersHard reject triggers
RelevanceDoes the option solve the assigned problemAnswers a different question, ignores a mandatory constraint, introduces an unrequested workforce or customer outcome, claims authority the exercise does not have
EvidenceWhat carries each material statementInvented citations, unsupported claims, a superseded source, a missing controlling record, an inference presented as established fact
AudienceCan the intended recipient use it correctlyConceals uncertainty, uses internal shorthand with customers, buries a required action, implies AI made a judgement reserved for a person
ConsequenceWhat happens if it is used or believedAn unauthorised promise, a recommendation beyond the drafter's remit, an unacknowledged material risk, a route that removes required human review

The need for hard triggers is visible in another peer-reviewed experiment. Dell'Acqua and colleagues studied 758 Boston Consulting Group knowledge workers on realistic consulting work. Across 18 tasks designed to be within GPT-4's capability frontier, AI assistance improved speed, completion and quality. On one complex task designed outside that frontier, AI-assisted participants were 19 percentage points less likely to reach the correct answer than the control group, even though their responses scored higher for coherence and persuasiveness (INFORMS journal article).

That outside-frontier result came from one task using a particular model and experimental design. It is not an error-rate forecast for your work. It demonstrates why fluency or persuasiveness cannot compensate for failed evidence, the same failure mode as a polished output that hid the expert safety work.

Use this prompt to prepare a draft REAC rubric before options are generated. The human task owner and relevant domain, risk, legal or compliance reviewer must set and lock the criteria, hard reject triggers and threshold before use.

Prompt
Draft a REAC rejection rubric for [TASK] using only the approved brief below.

For Relevance, Evidence, Audience and Consequence, propose:
- the task-specific question the criterion must answer
- observable evidence for scores 0, 1 and 2
- hard reject triggers that no score can offset
- the human role qualified to assess the criterion

Do not generate options, set decision rights, approve the task, assess compliance or make a regulated decision. Mark any missing criterion input [CHECK].

[PASTE APPROVED TASK BRIEF, SOURCE BOUNDARY, AUDIENCE AND HUMAN AUTHORITIES]

How do you record why an option survived?

Give each option an ID and preserve its submitted form. Then create an option-survival record with:

  • the task and rubric version applied
  • the human assessor and assessment date
  • each hard reject result with the supporting reason
  • REAC scores and evidence references for options without a hard reject
  • unresolved assumptions, source checks and specialist reviews
  • the survival decision: reject, hold for missing evidence or advance for further work
  • the authorised person who decides the next stage

Do not ask the generating model to be the sole judge of its own output. AI can organise options against locked fields, quote the supplied evidence and identify blanks. A qualified person verifies every classification and records the survival decision.

The CHI 2025 study by Lee and colleagues offers a useful caution, within its limits. Researchers surveyed 319 knowledge workers who already used generative AI at work at least weekly, and collected 936 first-hand examples of that use. Higher confidence in AI was associated with less self-reported critical-thinking effort, while higher task-specific self-confidence was associated with more. Participants described critical thinking shifting towards information verification, response integration and task stewardship (Microsoft Research).

This was a self-reported survey and does not establish that AI caused a change in critical-thinking skill. The practical implication is to cue the human work explicitly. A blank survival record is harder to wave through than an attractive paragraph.

Use this prompt to structure options against an already approved rubric. The named human assessors must verify the evidence, apply every hard reject and decide whether an option advances.

Prompt
Prepare draft option-survival records for [TASK] against REAC rubric version [RUBRIC_VERSION].

For each option:
1. quote the text relevant to each hard reject trigger
2. list the supplied evidence record supporting each material claim
3. propose REAC scores with reasons
4. identify unresolved assumptions and required human reviews
5. leave the survival decision blank for [HUMAN_ASSESSOR]

Use only the locked rubric, options and approved evidence. Do not repair an option while grading it, waive a hard reject, select a preferred option, verify a source not supplied or make the underlying decision. Mark missing support [CHECK].

[PASTE LOCKED RUBRIC, OPTIONS AND APPROVED EVIDENCE]

Consider a fictional example. [CUSTOMER_REMEDIATION_TEAM] uses an approved enterprise AI tool or equivalent to generate four draft explanations for customers affected by [FICTIONAL_SERVICE_ERROR]. Before generation, [HUMAN_TASK_OWNER] locks a REAC rubric. Every draft must answer what happened, state only facts from [APPROVED_INCIDENT_RECORD], explain the customer's required action in plain language and avoid promises beyond [AUTHORISED_REMEDIATION_DECISION].

Option A sounds reassuring but says all affected accounts have been corrected. The record confirms only that review is under way. It receives an Evidence hard reject.

Option B stays within the record but uses two internal acronyms and leaves the required customer action in the final paragraph. It fails a pre-set Audience hard trigger.

Option C states the confirmed facts, names the uncertainty, puts the action first and does not imply a remediation outcome. It has no hard reject and scores seven of eight, above the pre-set threshold. It advances to source verification, accessibility checking and the required legal, compliance and authorised communication reviews.

Option D tells customers no action is needed, although that has not been decided. It receives Evidence and Consequence hard rejects. The record retains why it was culled.

Option C has survived a rejection pass. It has not been approved for release. [AUTHORISED_COMMUNICATION_OWNER] remains responsible for the final message and its consequences.

Do this Monday

  1. Choose one option-heavy task. Use a bounded drafting activity where the team routinely generates several alternatives and a named human owns the next stage.
  2. Lock REAC before generation. Define task-specific Relevance, Evidence, Audience and Consequence tests, hard reject triggers, scoring anchors and a survival threshold with the relevant specialists.
  3. Generate a small set. Use an approved enterprise AI tool or equivalent and authorised inputs. Give every option an ID without silently combining or repairing them.
  4. Reject before scoring. Apply hard triggers first. Score only survivors, record evidence and leave unresolved questions visible.
  5. Advance, do not approve. Send the surviving option or options into the existing source, risk, legal, compliance and authorised decision process. Keep the survival records with the work.

Bottom line

Cheap drafting makes disciplined rejection more valuable. Lock Relevance, Evidence, Audience and Consequence before alternatives arrive, and let hard failures end an option before style wins support. Record why each survivor deserves further attention, then run the source checks and specialist reviews the work requires. The rubric narrows the field; an authorised person still owns every decision and consequence.

This article is general information and education only. It is not legal, compliance, financial or professional advice. Obligations vary by organisation and circumstance. Verify current requirements against the primary sources cited and seek advice specific to your situation.

References

  1. Shakked Noy and Whitney Zhang, "Experimental evidence on the productivity effects of generative artificial intelligence", Science, volume 381, issue 6654, pages 187-192. https://doi.org/10.1126/science.adh2586 (published 14 July 2023)
  2. Fabrizio Dell'Acqua and colleagues, "The Cybernetic Teammate: A Field Experiment on Generative AI and Teamwork", Organization Science, volume 37, issue 4, published online 12 June 2026, pages 1217-1242. https://pubsonline.informs.org/doi/10.1287/orsc.2025.20702
  3. Fabrizio Dell'Acqua and colleagues, "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality", Organization Science, volume 37, issue 2, published online 11 March 2026, pages 403-423. https://pubsonline.informs.org/doi/10.1287/orsc.2025.21838
  4. Hao-Ping Lee and colleagues, "The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers", Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, April 2025, pages 1-22. https://doi.org/10.1145/3706598.3713778

TheAICommand. Intelligence, At Your Command.

Frequently asked questions

What does REAC stand for?
Relevance, Evidence, Audience and Consequence. Relevance asks whether the option solves the assigned problem. Evidence asks what carries each material statement. Audience asks whether the intended recipient can use the option correctly. Consequence asks what happens if the option is used or believed. Each criterion carries hard reject triggers that no score elsewhere can offset.
Why reject before scoring rather than just ranking the options?
Because averaging lets a strong presentation carry a fatal flaw. A hard reject is a failure that cannot be offset, so it has to be applied before any total is calculated. Ranking first invites the most fluent option to accumulate support while its evidence gap stays unexamined.
Does a high score mean the option is approved?
No. Survival is not approval. A surviving option advances to source verification, accessibility checking and whatever legal, compliance and authorised communication reviews the work requires. The rubric narrows the field; an authorised person still owns the decision and its consequences.
Can the model that generated the options also grade them?
Not as the sole judge. AI can organise options against locked fields, quote the supplied evidence and identify blanks. A qualified person verifies every classification and records the survival decision. The prompt should also forbid the model repairing an option while grading it, because a quietly repaired option hides the very defect the gate exists to catch.
Does the research prove AI-assisted work is less accurate?
No, and the article does not claim it. In the Boston Consulting Group field experiment, AI assistance improved speed, completion and quality across 18 tasks inside the model's capability frontier. The 19 percentage point drop came from a single task deliberately designed to sit outside it. That is a demonstration of why fluency cannot substitute for evidence, not an error-rate forecast for your work.
LeadershipAI-assisted DraftingEditorial JudgementHuman ReviewFinancial Services
← Back to Leadership