Do not manage AI enthusiasm and AI refusal as personality problems. Give both people matched work under AI-required, AI-optional and AI-prohibited conditions, then assess quality, verification and boundary judgement against one visible standard.*
The employee who uses AI for everything and the employee who avoids it can create the same management problem: you do not know whether their method fits the task. Preference is not performance. Confidence is not control.
Test both against the work. Use three matched tasks, one requiring an approved AI tool, one making it optional and one prohibiting it. Score the outputs against the same quality, evidence, risk and explanation rubric. The pattern helps you choose the next evidence and whether the response is coaching, workflow repair, another observed task or, for an established concern, your organisation's ordinary human-led process.
This is a calibration exercise, not a trap. It uses prepared, fictional work samples and a disclosed rubric. It does not inspect private chats, count prompts, infer attitude from tool logs or let AI recommend employment action.
Why is preference the wrong thing to measure?
AI capability is uneven at task level. In a peer-reviewed experiment, 758 consultants were randomly assigned to complete realistic consulting work with no AI, GPT-4, or GPT-4 plus a prompt engineering overview. Across 18 tasks selected to sit within the model's tested capability frontier, participants with AI completed 12.2 per cent more tasks, worked 25.1 per cent faster and produced higher-quality work. On a separate complex task selected to be outside that frontier, participants using AI were 19 percentage points less likely to reach the correct solution (peer-reviewed field experiment).
That study does not establish a universal productivity rate. It tested one model, one occupational group and researcher-designed tasks. Its managerial value is the shape of the result: the same professional can benefit from AI on one task and be harmed by it on another that looks comparably difficult.
Enthusiasm does not guarantee scrutiny. A CHI 2025 study surveyed 319 knowledge workers about 936 first-hand examples of using generative AI at work. Higher confidence in AI was associated with less reported critical-thinking effort, while higher task-specific self-confidence was associated with more. Participants described critical thinking shifting towards verification, integration and stewardship (Microsoft Research). This was a self-report survey, not an experiment proving that AI reduced anyone's capability or that a confident user will underperform.
Refusal is not proof of poor judgement either. In five forecasting experiments, Dietvorst, Simmons and Massey found that people became more reluctant to use an algorithm after seeing it make mistakes, even when they had seen it outperform a human forecaster (Journal of Experimental Psychology: General, 2015). A later set of three incentivised forecasting studies found that participants were more willing to use an imperfect algorithm when allowed to adjust its forecast, even slightly, and performed better as a result (Management Science).
Those experiments involved statistical forecasts, not generative AI in Australian workplaces. They do not tell you why a particular employee declines an enterprise AI tool. They do show why "resistant" is an inadequate diagnosis. The person may be reacting to a visible error, unclear permission, poor task fit, loss of control or a genuine skill gap. Test the work before choosing the explanation.
How does the three-condition calibration work?
Build three comparable scenarios from one task family. Do not reuse the same scenario, because the first attempt teaches the answer. Match the complexity, source volume, time allowance and decision consequence as closely as practical.
AI-required. The employee must use the approved enterprise tool or equivalent within stated boundaries. This tests whether they can brief it, inspect its output, verify material claims, preserve uncertainty and take responsibility for the final work. It is not a speed contest.
AI-optional. The employee chooses the method before starting and records one sentence explaining why. This tests selection judgement. Choosing AI is not rewarded, and declining it is not penalised. The question is whether the chosen method suits the task and produces defensible work.
AI-prohibited. The employee completes a matched task without AI. This tests independent domain capability and the ability to respect a boundary. The reason may be that the exercise is checking unaided knowledge, that the approved tool is not suitable for the information, or that the workflow deliberately reserves a step for human analysis.
Keep one rubric across all three conditions:
- Quality: Is the answer accurate, complete and usable for the stated audience?
- Evidence: Can each material claim be traced to the supplied source, with conflicts identified?
- Risk: Were data, tool and human-decision boundaries followed?
- Judgement: Were uncertainty, exceptions and escalation points handled explicitly?
- Explanation: Can the employee explain the method, corrections and final recommendation in their own words?
Score method compliance separately from output quality. A polished answer produced in breach of the prohibited condition fails the boundary test. A sound manual answer in the AI-required condition may show domain strength, but it still leaves an adoption or instruction issue to understand. Knowing who can stop the model, and when, is part of the same boundary discipline.

Use this prompt to create a matched calibration pack from approved materials. A domain expert must review task equivalence, source accuracy, data handling and the final rubric before any employee sees it.
The calibration should be scheduled and explained like any other development activity. Tell participants what is being tested, which tools are approved, what will be retained, who will assess the work and how feedback will be used. Assess the submitted artefact and the review conversation. Do not turn continuous telemetry into a proxy for judgement.
What do the patterns actually tell you?
A single score does not tell you whether you have an enthusiasm problem, a refusal problem or no AI problem at all. Read the pattern across conditions.
Here is a fictional example using merge-field placeholders. In a synthetic operational-risk exercise, [TEAMMEMBERA] produced fluent work in the AI-required condition but could not trace two material claims and missed a stop trigger. Their AI-prohibited answer also missed the trigger. The pattern supports domain and verification coaching, not a verdict about excessive enthusiasm. Structured, supervised practice of the kind an AI apprenticeship builds is the natural next step.
[TEAMMEMBERB] produced accurate unaided work and selected an appropriate manual method in the optional condition. In the AI-required condition, they did not use the approved tool. The review found that [TEAMMEMBERB] believed the tool was prohibited for all risk material, based on an outdated local instruction. The first response is to correct the workflow and provide bounded practice. If the requirement later remains clear, reasonable and supported, repeated non-compliance becomes a separate human management question.
Use this prompt to organise de-identified results after the human assessor has scored them. The manager must test every suggested explanation with the employee and relevant subject-matter or people advisers. AI must not select coaching, warnings or formal action.
If an observed gap may become a performance matter, leave the calibration frame and use your normal process. The Fair Work Ombudsman's current best-practice guide recommends clear expectations, specific evidence, an opportunity for the employee to respond, agreed support, a reasonable period to improve and follow-up. It also notes that poor performance can have several causes, including unclear standards and missing knowledge or skills (Fair Work Ombudsman). Applicable policies, contracts, awards and enterprise agreements may add requirements, and the FWC's reading of AI use in a case shows how conduct with AI is already being tested formally. Obtain situation-specific advice before formal action.
Do this Monday
- Choose one task family, not one person. Select recurring, reviewable work such as a synthetic control assessment, policy comparison or operational-risk brief. Do not start with a disputed employee matter or live customer decision.
- Define the three conditions. Create matched AI-required, AI-optional and AI-prohibited scenarios. State why each condition exists and confirm the approved tool and information boundary.
- Publish one rubric. Give participants the five criteria, time allowance, assessor and feedback process before the exercise. Make clear that quality and method compliance are separate results.
- Run a small voluntary development pilot. Use fictional or de-identified material. Review whether tasks were genuinely comparable and whether several people hit the same confusing instruction.
- Hold a human review conversation. Ask the employee to explain the method, corrections and uncertainties. Agree the next response: coaching, clearer instructions, workflow repair, another observed task or referral into the ordinary performance process with appropriate advice.
Bottom line
The AI enthusiast does not earn trust by using the tool, and the AI refuser does not fail by declining it. Both earn trust by matching method to task, producing defensible work and respecting the boundary in front of them. Test those behaviours under required, optional and prohibited conditions. Manage the observed pattern, not the label.
This article is general information and education only. It is not legal, compliance, financial or professional advice. Obligations vary by organisation and circumstance. Verify current requirements against the primary sources cited and seek advice specific to your situation.
References
- Fabrizio Dell'Acqua and colleagues, "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality", peer-reviewed article published online 11 March 2026. https://pubsonline.informs.org/doi/10.1287/orsc.2025.21838
- Lee and colleagues, "The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers", CHI 2025. https://www.microsoft.com/en-us/research/publication/the-impact-of-generative-ai-on-critical-thinking-self-reported-reductions-in-cognitive-effort-and-confidence-effects-from-a-survey-of-knowledge-workers/
- Dietvorst, B. J., Simmons, J. P. and Massey, C., "Algorithm Aversion: People Erroneously Avoid Algorithms After Seeing Them Err", Journal of Experimental Psychology: General, 2015. https://doi.org/10.1037/xge0000033
- Dietvorst, B. J., Simmons, J. P. and Massey, C., "Overcoming Algorithm Aversion: People Will Use Imperfect Algorithms If They Can (Even Slightly) Modify Them", Management Science, 2018. https://pubsonline.informs.org/doi/10.1287/mnsc.2016.2643
- Fair Work Ombudsman, "Managing underperformance best practice guide", content last updated 16 January 2026. https://www.fairwork.gov.au/tools-and-resources/best-practice-guides/managing-underperformance
TheAICommand. Intelligence, At Your Command.


