Most AI rollout debates are about access. A field experiment suggests the two decisions that actually shape the result are whether the assistant is grounded in your own knowledge base, and who receives it first. Those choices produce different organisations three months later.
Almost every study a leader has been handed about AI at work measures what one person produces. Faster drafting, more code, better analysis. Useful, and incomplete, because the thing a leader manages is not a pile of individual outputs. It is a network of people who ask each other for help.
A randomised field experiment out of the University of Mannheim and INSEAD measured that instead. It asked what a grounded AI assistant does to who talks to whom, and who gets asked for knowledge. The results give a rollout decision an evidence base it has not had, provided the caveats travel with it.
What did the experiment actually test?
The design is unusually good for this subject. The researchers studied 316 employees across 42 teams at a technology services firm in Central Europe, randomly assigning teams to receive access to a grounded assistant tailored to the firm's knowledge base, or to continue as usual without access to AI. The draw put 25 teams in the treatment group and 17 in the control group. Some employees were generalists working in sales, others were specialists in technical support.
Outcomes were measured over a two-week period before the treatment and again three months after. Standard errors were clustered at the team level across 42 clusters, with analysis at the individual level and findings qualitatively replicated on team-level aggregated data.
Two design choices deserve a leader's attention. Randomisation happened at team level, which is how the intervention would actually land in an organisation. And the follow-up sits three months out, not at the end of a single session, so the measurement reaches past the novelty of a first sitting. Three months is also where the paper stops, and the authors say so: the window captures the initial effects and cannot show whether the changes hold.
The word doing the heavy lifting is grounded. The treatment was an assistant customised with organisation-specific knowledge, not a generic chatbot. The authors theorise that a tool customised to local context, through fine-tuning or retrieval-augmented generation, can change a person's centrality in the collaboration network by making them easier to work with, and their centrality in the knowledge-sharing network by changing their value as a source of knowledge. The tool itself was built with retrieval-augmented generation on GPT-4 and later GPT-4o, and it reached the treated teams in September 2024, which dates the evidence. A generic assistant was not tested, so nothing here says a plain chatbot does the same thing.

What actually changed in the network?
Both networks expanded, and not marginally. Degree centrality in the collaboration network increased by 7.77 in the treatment group compared with 1.12 in the control group, a difference of 6.65 at p below .001, with similar patterns for in-degree and out-degree. In the knowledge network, degree centrality increased by 5.21 in the treatment group compared with 0.84 in control, a difference of 4.37 at the same significance level.
That is the headline worth carrying into a leadership team meeting. The intuition that an assistant makes people more self-sufficient, and therefore quieter, did not show up. Connections went up. The authors describe the effect in their setting as predominantly augmentation-oriented, saying the AI did not replace human interaction or work but amplified it, and they confine that observation to their context.
Worth naming a tension rather than smoothing it over. This site has covered evidence that AI narrows the pool of ideas a team generates. Those findings are not in conflict, because they measure different things: convergence in the content people produce, against expansion in the interactions people have. A team can talk to more colleagues and still arrive at more similar answers. A leader who only reads one of those results will manage for the wrong risk.
Where does the role difference show up, and where does it not?
Here is where careless summaries will go wrong, so it is worth being exact. One scope note first: the role comparisons run on the 282 non-managers in the sample rather than the full 316, and across 41 clusters rather than 42, because the paper excludes the 34 managers from that analysis.
Specialists did not become more connected than generalists. Overall degree centrality rose 7.34 for specialists and 8.71 for generalists, with in-degree up 3.26 against 4.45 and out-degree up 4.08 against 4.26. The interaction term testing whether AI affected collaboration centrality differently by role was not statistically significant, and the authors state plainly that the hypothesis was not supported.
Among the six network measures the difference appears in one only, and it is the one that matters for knowledge. For in-degree centrality, the measure most directly reflecting standing as a source of knowledge, specialists gained 6.26 while generalists gained 2.28, with a negative and significant interaction term. The authors' reading is that specialists became disproportionately more central as knowledge providers.
That is a specific and useful claim. Give a grounded assistant to your deepest experts and, in this setting, more colleagues came to them. The plausible mechanism is not that the tool made the expert smarter, but that it lowered the cost of turning their expertise into something a colleague could use, which is a packaging problem most organisations have rather than a knowledge problem.

What happened to output?
Both groups improved. Generalists moved from a baseline mean of 12.7 projects to 16.30 after implementation. Specialists moved from 4.29 to 5.42. Both changes were significant.
The comparison needs care, because it is easy to flatten into a slogan. In percentage terms the two are close: 28.3% for generalists against 26.4% for specialists. What differs is the absolute gain, which was substantially larger for generalists because they work from a much bigger baseline. So the honest statement is not that generalists benefit more from AI. It is that in this firm, the same proportional lift applied to a larger volume of work produced more additional throughput.
For a leader deciding where to start, that distinction changes the reasoning. If the goal is total volume this quarter, the arithmetic favours whoever handles the most work. If the goal is to raise the capability of the organisation, the arithmetic is not the right guide at all.

What does the sequencing decision look like in practice?
Take a division with two populations: [SPECIALIST_TEAM], a small group holding deep technical knowledge that everyone else eventually needs, and [DELIVERY_TEAM], a much larger group carrying the bulk of client volume. There are enough licences for one group this quarter.
The throughput case is easy to build and easy to defend. [DELIVERY_TEAM] handles the most work, so the same proportional lift returns the most additional output, and the number lands inside the reporting period. The diffusion case is slower and harder to put in a paper. Give it to [SPECIALIST_TEAM] first, ground it in their documentation and prior work, and what you are buying is a reduction in the cost of getting their knowledge out of their heads and into a form a colleague can use.
Neither is wrong. The failure is choosing by default, which usually means choosing throughput because it is the case that fits a quarterly template. So make the trade explicit. If [DIVISION_LEAD] cannot say which of the two this rollout is for, the pilot list is being drawn up by whoever asked first, and the result three months later will be whatever that produced.
One caution on the diffusion path. If it works, a small group becomes the place more questions land, and that load is real work that arrives without a plan.
What does this evidence not support?
Five limits, four of them named by the authors, and all load-bearing.
The outcomes are self-reported. Performance data was self-reported rather than extracted directly, to respect personal data protection guidelines. So the correct statement is that specialists reported being asked more often, not that they were measurably asked more often. Self-report and system measurement diverge in predictable ways, and the direction here is unknown.
Role was not randomised. The authors are explicit that while AI adoption was randomly assigned, being a generalist or a specialist was not, so the role comparisons cannot be read as causal estimates of heterogeneous treatment effects at all, and are instead correlational patterns describing how the two roles differ as they already exist in that organisation. Read those results as a strong signal about where to look, not as proof of what causes what.
It is one firm, in one sector, in one region. Nothing establishes that a professional services firm, a bank or a claims operation behaves the same way.
Three months is the whole horizon. The authors name that as a limit of their own: the window captures the initial effects, and they cannot say whether the network changes persist, are amplified, or reverse over a longer period as people and the technology settle around each other. Nothing here shows the effect is durable.
And it is a working paper, not yet peer reviewed. The authors also claim this is among the first field experiments to show an AI tool reconfiguring collaboration patterns measurably, which is their characterisation rather than an established fact.
That is a solid piece of evidence with real limits, which is the normal condition of useful management research. It is strong enough to shape a sequencing decision. It is not strong enough to justify a restructure.
Do this Monday
- Write down the outcome before the pilot list. Decide whether this rollout is for throughput this quarter or for spreading expertise across the organisation. The two point at different first cohorts.
- Check whether your assistant is actually grounded. The measured effects attach to a tool customised on the firm's own knowledge. If your deployment is a generic licence with no connection to your material, this evidence does not describe it.
- Pick the first cohort deliberately. For diffusion, start with the specialists whose knowledge others already need and cannot easily get. For volume, start where the work is heaviest.
- Measure one network question, not just usage. Add a single item to an existing pulse survey asking who people went to for help this month. Usage counts will not show you the effect this study found.
- Baseline before, not after. The study measured two weeks before and three months after. Without a before, the interesting result is unrecoverable.
- Protect the newly central expert. If more colleagues start routing questions to a small group, that is the intended effect creating a real load. Name it, resource it, and check it against their other commitments.
Bottom line
The rollout decisions that matter are not how many licences to buy. They are whether to ground the assistant in your own knowledge base, and who receives it first. In this experiment, a grounded assistant expanded both collaboration and knowledge networks well beyond the control group, made specialists markedly more central as sources of knowledge, and delivered the largest absolute output gain to generalists working from a bigger baseline. Collaboration centrality itself did not differ by role, so the case for starting with specialists rests on knowledge diffusion, not on connection. Choose the outcome first, then the order, then measure the network rather than the usage. And carry the caveats honestly: self-reported outcomes, one firm, roles that were not randomised, a three-month horizon the authors say cannot show whether the changes hold, and a paper still awaiting peer review. Related reading: why the advantage cannot sit with one team, who owns the shared skill library, and what happens to the idea pool.
This article is general information and education only. It is not legal, compliance, financial or professional advice. Obligations vary by organisation and circumstance. Verify current requirements against the primary sources cited and seek advice specific to your situation.
References
- Ralf Buechsenschuss, Irmela Koch-Bayram, Torsten Biemann and Phanish Puranam, GenAI Adoption Increases the Density of Knowledge and Collaboration Networks: Evidence from a Field Experiment, INSEAD Working Paper No. 2026/22/STR, revised version of 2026/01/STR, dated 1 April 2026 (working paper, under review, not peer reviewed). https://doi.org/10.2139/ssrn.6028034
- INSEAD, working paper record 2026/22/STR, full text of GenAI Adoption Increases the Density of Knowledge and Collaboration Networks, accessed 21 September 2026. https://sites.insead.edu/facultyresearch/research/doc.cfm?did=75355
TheAICommand. Intelligence, At Your Command.


