The document stopped telling you who did the thinking.
Four numbers carry the argument. Among participants who all had an AI assistant and all leaned on it heavily, the graded deliverables scored 1.878 and 1.972 standard deviations, a gap the authors could not distinguish from zero at p = 0.206. When the assistant was taken away and the same people were asked to explain the root cause of the problem in their own words, those two groups scored 0.732 and 0.065, a difference significant below the 1 percent level. Same tool, same reliance, indistinguishable work product, two entirely different answers to what the person could do without it.
One paragraph of setting before anything gets built on that. The source is "Does generative AI narrow education-based productivity gaps? Evidence from a randomized experiment", by Guillermo Cruces, Diego Fernandez Meijide, Sebastian Galiani, Ramiro H. Gálvez and María Lombardi, circulated as NBER Working Paper No. 34851 in February 2026, revised May 2026, and posted to arXiv as 2608.04198 on 4 August 2026. It ran online in Argentina between September and November 2025 with 1,174 adults aged 25 to 45, deliberately outside firms, and has not been peer reviewed. Australian professional work is not the setting, and what a leader takes from it is an extension rather than a finding.
What the experiment actually measured
Participants were sent a report with a figure and a table, asked to diagnose the root cause of a business problem, propose a solution and reply by email. The experiment was designed for twenty minutes, mean completion was twenty-one, and prize money was tiered on relative performance so the incentive was real. The treatment group got a custom assistant preloaded with the same report, every interaction logged. Responses were graded against task-specific rubrics built from pre-registered criteria, by an AI-assisted procedure validated against two human graders, with correlations above 0.9.
The paper's own headline is education. In the control group, higher-education participants outperformed lower-education participants by 0.548 standard deviations; with the assistant that gap fell to 0.139, closing about three quarters of it. That result earns the study its credibility. It is not the finding leaders should act on.
The one that matters here sits in Section 5.2, where the authors split the treated group two ways: by how much of the task they asked the assistant for help with, and by whether they spent more or less time on it than the treatment-group median.
The third cell is worth its own beat. Participants with high engagement and low AI assistance also did well on the unassisted follow-up. Across the four groups, what predicted the deliverable was how much AI help you took, and what predicted the unassisted answer was engagement. As the authors put it, exposure to a correct AI-generated answer is not by itself sufficient for follow-up performance.
Three limits travel with that table and none belong in a footnote. The 2x2 is exploratory, not pre-registered and not causal: randomisation was into AI access, not engagement, so people sorted themselves into these cells by their own behaviour. Engagement is a proxy, time on task against the median, not a measure of cognitive effort. Controls for education, age, gender, employment status and work experience leave the pattern quantitatively similar, which is reassuring about composition but does not make an observational comparison causal.

This is not a deskilling study, and saying so matters
The easiest way to misreport this paper is to file it under AI making people worse. It found close to the opposite. Treated participants did not perform worse than controls once AI was removed, and AI access modestly improved unassisted follow-up performance among lower-education participants, by about 0.171 standard deviations and significant at the 5 percent level. The higher-education estimate of 0.071 was small and not significant.
That cuts against the frame in protect the reps, which is about deciding on purpose which recurring work stays human so capability does not erode over years. Both can hold. That piece is about which reps to keep. This one is about not being able to tell who did the rep.
The authors are careful about what the follow-up measures and this article should be too. They use carry-over rather than learning, because the design sets no new unassisted task drawing on the same skills and the follow-up runs immediately afterwards. It captures short-run unassisted performance after AI exposure, not durable skill formation. They are equally explicit that the results are specific to one self-contained task and to one generation of AI capability.
The signal problem, which is ours rather than theirs
Here is the reading, and it is a reading. What changed for a leader is not that people got worse. It is that the artefact has gone quiet on a question it used to answer.
Flag the inference plainly. The study established that rubric-graded quality was statistically indistinguishable between the two heavy-use groups. It did not test manager perception, and the graders did not know the participants or the shape of their previous work. That a manager reading the document cannot tell the two apart is our inference from the rubric result, not something the researchers measured.
Even as an inference it is uncomfortable, because the deliverable is the evidence most leadership systems run on. Stop counting prompts argues for moving off activity metrics onto the work that actually improved. This puts a number on the outcome metric going quiet too. It still measures the output. It no longer measures the person.
Nor is it the review problem. The review tax is about the cost of checking whether AI output is correct. Here the disengaged group's output was correct, and nothing needed catching. The issue is what a correct document fails to tell you.

The operating move: ask before you read
The study hands you an instrument, and it is smaller than a policy. To measure retention, the researchers asked participants, with the tool switched off, to summarise the root cause of the problem, framed as a scenario in which their boss, who had not yet read their response email, asked about it directly.
That is a research instrument. The authors never tested it as a management practice and make no claim that asking changes behaviour or builds anything. What follows is our adaptation of it, worth trying rather than validated.
- Pick work that has already landed. Finished and sent, not in flight. You are not helping with the task, you are talking briefly about one that is done.
- Ask before you open the document. The framing is the study's own. Say plainly that you have not read it yet, then ask for the root cause and the main trade-off they made.
- Two minutes, spoken, no screen. The moment notes appear you are reading the artefact again through a different window.
- Listen for structure, not fluency. Can the person name the mechanism, the option they rejected, and why? A confident restatement of the recommendation sounds different from an explanation of what drives it.
- Say out loud what this is and is not. Not a test, not scored, not going anywhere. Say it the first time and mean it, because the practice works only while people believe it.
- Run it on a rotation, not on suspicion. Everyone, in turn. Aimed at individuals it becomes an accusation, and the point is that the artefact gave you no grounds to aim at anyone.
- Act by changing the work, not the rating. If someone cannot narrate the reasoning behind good output, change how the next piece is scoped, paired or reviewed.
The judgement boundary
Four lines, and the first matters more than the rest.
Do not turn this into time tracking. Time on task was a research proxy inside a twenty-minute experiment, chosen because it sat in a log. It was never a management metric, and a leader who responds by monitoring how long people spend on documents has read the study backwards. In an Australian setting it also walks into psychosocial hazard and privacy territory that this practice has no reason to open.
Do not make it an assessment. Framed as a test, a gate or a scored input, it teaches people to prepare an answer, which destroys the only thing it was measuring.
Do not extend it to calibration, promotion or hiring. Any move that way is ours, not the paper's. The authors make no claim about any of them and expressly decline to predict wage effects. Treat it as an implication to argue about, not a finding to act on.
Do not scope it to juniors. The apprenticeship argument is about what new starters are owed and stands on its own. This is different: the engagement split is education-independent and held under controls, so the person most likely to produce an excellent document they cannot narrate is not reliably the newest.
A worked example
[TEAM_MEMBER] sends [DELIVERABLE], a vendor comparison with a clear recommendation. It is well argued, the criteria are sensible and the pricing holds up.
Under the old habit you read it, agree, and move to implementation. Under the new one you call before opening it: what is actually driving the recommendation, and what did you give up to get there?
One answer names the mechanism, that two of the four criteria measure much the same thing so the scoring overweights integration depth, and that the runner-up was dropped on a timeline the vendor has since revised. Another restates the recommendation confidently and cannot get behind it.
Both documents are good and nobody has done anything wrong. What differs is what you now know, and it took two minutes. In the second case the response is to the work: pair the next comparison, or ask for the rejected option written up beside the chosen one. This is a cousin of the polished output that hid the expert, except that piece is about judgement invisible inside one artefact and this is about two that are indistinguishable.
Bottom line
In a randomised experiment run in Argentina in late 2025, participants who all used an AI assistant heavily produced deliverables that graded the same whether or not they engaged with the problem, 1.878 against 1.972 standard deviations at p = 0.206. Take the tool away, ask them to explain the root cause, and the same two groups scored 0.732 and 0.065. The paper found no average harm from AI access and a small positive carry-over for lower-education participants, so this is not deskilling evidence. It is evidence that the work product has stopped carrying something a leader used to read off it, in one task, in one setting, in an exploratory analysis the authors do not call causal. The fitting response is small: ask for the reasoning before you read the document, on work that has landed, as a conversation and nothing else.
Do this Monday
- Pick three pieces of finished work from last week. Anything where the output was good and the reasoning was never discussed.
- Ask one person, before you read theirs. Root cause and trade-off, in their own words, two minutes.
- Say the boundary out loud, before the question. Not a test, not scored, not going anywhere.
- Write down what you could not have known from the document. That gap is the finding, in your own team, for the price of one conversation.
- Put it on a rotation. Everyone, in turn. If it only happens to the people you already wonder about, it is a different practice and a worse one.
The productivity is real and none of this asks you to give it back. But leaders read work products to learn about people as well as about work, and one of those signals has gone flat. The document still tells you whether the job got done. Ask if you want to know anything else.
References
- Guillermo Cruces, Diego Fernandez Meijide, Sebastian Galiani, Ramiro H. Gálvez and María Lombardi, "Does generative AI narrow education-based productivity gaps? Evidence from a randomized experiment", NBER Working Paper No. 34851, February 2026, revised May 2026. nber.org/papers/w34851
- The same paper as an arXiv preprint, 2608.04198, posted 4 August 2026. arxiv.org/abs/2608.04198
- Pre-registration, AEA RCT Registry trial 0016607. socialscienceregistry.org/trials/16607
TheAICommand. Intelligence, At Your Command.


