They can see the tool. They cannot see the check.
The standard advice about leading AI adoption is to use it yourself, visibly, because a manager's own behaviour is the strongest signal a team gets about what is permitted and what is worth learning. That advice holds up. What it does not account for is what happens two steps later, and a paper posted to arXiv on 20 August 2026 gives that second step a name.
The mechanism
Ahana Biswas models AI reliance as a population process rather than an individual one. Agents repeatedly face a task and choose between solving it alone, accepting an AI answer, or verifying it, updating a belief about AI quality as they go and, when connected, learning from peers.
Four results come out of it, and the order matters.
The environment sets the baseline. Task difficulty and AI quality between them fix how much overreliance you get before anyone looks at anyone else.
Social learning, on its own, does not cause overreliance. The paper is careful about this: connectivity "moves the aggregate only when influence transmits beliefs". A better connected team is not automatically a more credulous one. This is the result most people would have guessed wrongly.
Social proof is the different thing. In the paper's words, "visible unverified use suppresses verification and tips the population into collective overreliance". Not shared beliefs about the tool. Observed behaviour with it.
And the fourth result is the useful one. "Feedback design can prevent collapse: making verification visible or dampening social proof reverses it."

Say plainly what this is. It is an agent-based model, so the agents are simulated and the findings describe the dynamics of a model rather than a measured workplace. Nobody has run it on a real team, and the paper agrees: its own stated next step is to bring the model to data and estimate a few key quantities from longitudinal traces. It has been accepted for CSS 2026, the Computational Social Science Society of the Americas conference in late October, so it has cleared conference review rather than sitting as an unreviewed preprint. What it offers is a mechanism that explains something most leaders have already watched happen, and a corrective cheap enough that it does not need to be proven before it is worth doing.
Why the advice was half right
Go back to the original move. Setting your team's AI norm by using it visibly works because adoption is a permission problem before it is a capability problem. People need to see that the tool is allowed, that it is not cheating, and that someone senior is prepared to be seen learning it.
All of that is true. The trouble is what a team can actually observe.
Your use is observable. It is in the meeting, in the shared document, in the sentence about which tool you asked. Your checking is not. It happened at your desk, in your head, in a tab nobody saw, in the ninety seconds where you read the source and confirmed the number. Nothing about it is visible unless you say it out loud.
So a team that copies you faithfully copies the observable half. Not because anyone is lazy, and not because the standard was not communicated. Because that was the part on display.
The paper's contribution is to say what happens when that runs across a group rather than one person. Each unverified use that others can see makes verification look less necessary, which produces more unverified use, which makes it look less necessary still. It compounds, quietly, and it does not announce itself until something goes out the door wrong.
Written out, the gap between what happened and what was transmitted is uncomfortable to look at.
Every row in the second column is accurate. That is the problem. Nobody was misled, and the signal still came out wrong, because half the process was never on display.
What this is not
Two boundaries, because both are easy to blur.
This is not the review problem. The AI review tax is about time saved leaking back out as checking and correcting, and how to design the checking so the gain survives. That assumes checking is happening. This piece is about a group in which it has quietly stopped.
The paper draws the same line from the other direction. Of the levers it tests, "merely reducing verification friction is the weakest lever", because "cheaper checking does not counter the social pull toward unverified use". Making checking cheaper and making it visible are different interventions, and on this mechanism only the second one moves the group.
It is also not the escalation problem. Silence is not human oversight is about someone who can see a suspect output and has no safe route to stop it. Here nobody is looking, so there is nothing to escalate. A challenge route in a team that has stopped verifying is a fire door in a building where nobody smells smoke.
How would you know?
The cascade is hard to see from inside because nothing goes wrong for a while. Output stays good, because the model is mostly right, which is exactly the condition under which checking feels least necessary and matters most.
The paper expects that experience. Its limitations section says the sharp version of the tipping is "a mean-field property", and that the agent-based model itself "shows a smooth crossover", with overreliance expected to "build gradually as social proof strengthens rather than to switch on at a sharp threshold". Verification erodes continuously, which the paper calls "both less dramatic and harder to notice in deployment". There is no alarm to wait for, which is why you have to go looking.
Four questions, none of which require a survey.
When did you last hear someone say what they checked? Not that they used AI. That they verified something, and what they verified it against. If you cannot recall an instance in the last month, the behaviour is not being observed by anyone.
What happens when someone finds an error in AI-assisted work? If the response is embarrassment, the team has learned that a caught error is a failure rather than a system working. Catching things is the output you want more of.
Can two people in your team describe checking the same way? Ask separately what verification means for a document like the one they last sent. Divergent answers are the point at which everyone starts defining it privately, and private definitions drift downward under time pressure.
Who is the most careful person in the room, and has anyone noticed? If the answer to the second half is no, the team's best example is invisible, which by this paper's mechanism means it is not an example at all.
The move
Make the checking the observable part. That is the whole intervention, and it costs a sentence.
One caveat carried straight from the paper, because it changes how you do this rather than whether you do it. The reversal is "parameter-dependent". It is "a counter-cascade above a threshold", not a guaranteed outcome, and below that threshold "the recovery is partial". Read practically, intensity matters. Saying it once, quietly, in a single covering message is not the setting the model reverses at. The visible checking has to be regular enough and public enough that it becomes the thing people expect to see.

- Narrate the check when you share the work. Not "I used AI to draft this". Say what you verified and against what. "The two figures are from the source document, I checked both, the third claim I cut because I could not find it." Twenty seconds, and it does the work an entire policy does not.
- Report the clean checks too. If you only mention verification when it caught something, you have taught the team that checking is what you do when you are suspicious. Reporting a check that found nothing wrong is what makes it routine rather than accusatory.
- Bring one caught error to the team, without blame attached. The error is not the point. The visible act of having looked is the point, and someone senior modelling it removes the cost of admitting a check happened.
- Change the review question. Ask "what did you check and how" before "is this good". The first question makes verification a normal thing to have a ready answer to. The second one rewards a polished artefact, which is exactly the signal that got you here.
- Name what a check consists of for your work. Sources traced, figures matched to a document, the claim you could not verify and cut. If nobody can say what checking means in your team, everyone will define it privately and downward.
- Watch the tail, not the average. The paper's cascade is a group-level effect. The question is not whether your team believes in verification, it is whether anyone has seen it lately.
TheAICommand works to the Verified Draft Method: de-identify the inputs, ground the model in your own source material, keep a person at the decision point, verify against the primary source, and log what happened. The fourth step is the one this paper is about, and the finding is that doing it silently is worth less than doing it out loud.
A worked example
[TEAM] of nine, six months into everyday AI use, no incidents, no complaints. Output is up and quality has not visibly moved.
The manager, [MANAGER_NAME], is a careful checker and always has been. Every AI-assisted document they circulate has been read against the source. None of that has ever been mentioned, because mentioning it felt like either boasting or hedging.
What the team has observed for six months is a senior person confidently sharing AI-assisted work. They have inferred the reasonable thing from it, which is that the tool is reliable enough to share from.
The change is one line in the covering message, said every time rather than once. "Checked the three numbers against the report, the second one was wrong in the first draft." Nothing else changes. No new policy, no new step, no additional review layer.
What it alters is what the team can copy. Within a few weeks the same line starts appearing in other people's messages, because it has become a thing colleagues do rather than a thing the policy requires. This is not a promise about your team, and the paper cannot support one. It is what the model says the mechanism would do, and it is a cheap thing to test on your own.
The uncomfortable part
If this is right, then the more competent and careful you are, the worse your silence is.
A leader who checks everything, thoroughly, and says nothing about it is generating maximum social proof for unverified use. Every document they circulate looks, from the outside, exactly like a document nobody checked. Their care is invisible and their confidence is not.
That is a strange thing to have to fix by talking more about your own process. Most people who are good at verification are good at it partly because they do not make a performance of it. The finding, if it holds, is that the performance is the part with the transmission effect, and the quiet competence stays where it is.
Do this Monday
Pick the next piece of assisted work you send to your team and add one sentence to the covering message naming what you checked and what you found. Not that you used a model, and not that you reviewed it, but the specific thing: which figure you traced to source, which clause you opened, what you changed as a result. Do it on three pieces this week. The point is not the disclosure. It is that the checking becomes observable, so it becomes the thing that gets copied.
Bottom line
Visible use is still the right way to start. It is just not the whole signal. A team copies what it can observe, and checking is the part that happens where nobody is looking, so a group can drift into collective overreliance without a single person deciding to stop caring. The modelling that names this mechanism is a simulation and should be held as one. But the corrective it points to costs one sentence in a covering message, works whether or not the model is right, and is the only version of a verification standard that gets copied rather than filed.
References
- Ahana Biswas, "Modeling AI Overreliance as a Complex Adaptive System", arXiv 2608.19616, submitted 20 August 2026, accepted at CSS 2026. arxiv.org/abs/2608.19616
TheAICommand. Intelligence, At Your Command.


