Most AI pilots are judged on the wrong timescale using the number that was easiest to export. That is not a failure of rigour. It is what happens when nobody decided the evaluation design before the pilot started, so the design got decided by whatever the tool's dashboard produced.
There is now a piece of evidence that makes the cost of that concrete.
What did a thirteen-month look actually find?
In "Impacts of Generative AI on Agile Teams' Productivity: A Multi-Case Longitudinal Study", published in the proceedings of the 2026 IEEE and ACM Third International Conference on AI Foundation Models and Software Engineering, Rafael Tomaz, Paloma Guenes, Allysson Allex Araújo, Maria Teresa Baldassarre and Marcos Kalinowski followed three agile teams at a large technology consulting firm for approximately thirteen months.
The method matters as much as the result. The authors combined quantitative telemetry drawn from Jira, SonarQube and Git with qualitative surveys, and compared sprints before adoption of internal GPT tools and GitHub Copilot with sprints after. That is a before-and-after design running across more than a year of real delivery work, rather than a controlled task run once with volunteers.
What they found was a sharp increase in performance and perceived efficiency concurrent with flat developer activity. Teams reported improved well-being alongside the performance gains. The authors read the pattern as generative AI improving the quality of the work rather than the volume of it, and concluded that multi-dimensional assessment frameworks such as SPACE are essential to capture the true impact, because a single-metric evaluation would miss these effects entirely.
Sit with the shape of that for a moment. Activity, the dimension almost every pilot reports on because it is the one the tooling emits without being asked, did not move. If you had run this pilot, measured commits or tickets closed, and reported to a steering committee, your honest finding would have been that nothing happened.

Why the default evaluation is stacked against you
Two habits combine badly.
The first is the short window. Pilots are scoped to six or eight weeks because that is how long a leader can hold attention and budget without a decision point. Within that window, a team is still learning the tool, still discovering which tasks it suits, and still absorbing the cost of changing how they work. The literature the authors position against is explicitly short-term and individual-focused, which is why they built a longitudinal, team-level study in the first place.
The second is the available metric. Every tool ships usage analytics. Almost none ship outcome analytics, because outcomes are specific to your work and usage is generic. So the number that arrives without effort is a count of activity, and the number that would answer the question requires somebody to define quality for your context and then go and measure it.
Put those together and the default pilot is a short observation of the one dimension least likely to move, which produces one of two bad outcomes. Either the leader concludes there is no effect and stops something that was working, or the leader overrides the data with enthusiasm and scales something on the strength of anecdote. Both are decisions made without evidence. Only one of them looks like it.
The operating move: pre-commit, in writing, before you start
This is a fifteen-minute exercise and it has to happen before the first licence is issued, because its whole value is that it cannot be adjusted once the results are visible.
Fix the window. State how long the pilot runs before anyone renders a verdict, and separate that from how long it runs. A reasonable pattern for knowledge work is a learning period where you deliberately do not judge, followed by an assessment period. Naming the learning period out loud also removes the pressure on the team to look productive while they are still working out where the tool helps.
Fix the dimensions. Choose three or four, and require at least one that is not a count. The same discipline applies to judging people rather than pilots, where speed is not the skill. The SPACE framing the authors point to spans satisfaction and well-being, performance, activity, communication and collaboration, and efficiency and flow. You do not need to adopt the framework wholesale. You need to avoid a scorecard where every row is a volume.
Fix the instruments. For each dimension, name where the number comes from and who pulls it. A dimension with no instrument is an intention. If the honest answer is that quality will be assessed by [MANAGER_NAME] reading a sample of ten pieces of work, write that down; it is a legitimate instrument and it is far better than a proxy nobody believes.
Fix the baseline. Capture the same measures for a period before the tool arrives. This is the step that gets skipped, and skipping it converts the entire exercise into an opinion. The study's design turns on exactly this comparison.
Fix the stopping rule. Write the result that would cause you to stop, and the result that would cause you to expand. A pre-committed stopping rule is the only reliable defence against the sunk cost of a tool everyone has already learned.
Where a leader has to stay in the loop
The pre-commitment is not a way to take yourself out of the decision. It is a way to make sure you are making the decision rather than ratifying whichever number arrived.
Two judgements stay yours throughout.
The first is what quality means for your work. No framework supplies this. In a claims team it might be the rate of decisions that survive reconsideration. In a policy team it might be the number of drafts a subject matter expert has to send back. In a client-facing team it might be the proportion of work that reaches the client without rework. Somebody has to say it out loud, and the person accountable for the output is the only one with standing to.
The second is the honest reading of a flat number. A dimension that does not move is information, not failure. In this study, flat activity alongside rising performance and perceived efficiency was the most interesting result in the paper. A leader who has pre-committed to multiple dimensions can say that. A leader who committed to one cannot.
There is also a limit to state plainly, because pretending otherwise would be exactly the error this article is about. This is three teams in one technology consulting firm, in software engineering, measured through developer tooling. The design lesson transfers. The specific pattern is evidence about software teams and should be described that way when you cite it internally.
The cost of a longer window
Extending the window is not free, and pretending it is will get you a different bad decision.
A thirteen-month observation of a real team imports thirteen months of everything else. People join and leave. The work itself changes. A reorganisation moves two of your best people. A regulatory deadline compresses one quarter and empties the next. A new system lands that has nothing to do with AI and changes cycle time anyway. By the time the signal is large enough to trust, the explanation for it is contested.
The answer is not to shorten the window back down. It is to keep a plain timeline alongside the measures.
Run a single dated log for the duration of the pilot, and record every event that a reasonable person could later point to as an alternative explanation: team composition changes, workload spikes, system changes, process changes, leave patterns, anything that moved. It takes a minute a fortnight. Its value arrives at the end, when somebody asks whether the improvement was the tool or the two experienced people who joined in March, and the honest answer is available rather than reconstructed.
This also protects the team. Without the log, a pilot that coincides with a hard quarter reads as a tool that did not work, and the people who used it carry that. With the log, the leader can say what else was happening and separate the two questions.
It is worth noting that the study's own design leans on this. A before-and-after comparison across sprints only means something if you can speak to what else changed between them, which is part of why a multi-case design across three teams is stronger evidence than one team observed for longer.
A worked example
[TEAM_NAME] is a twelve-person operations team about to trial an AI drafting assistant. The pre-commitment, agreed and circulated before licences are issued, reads roughly like this.
The pilot runs for six months. No verdict is formed in the first eight weeks, which are named as the learning period. Four dimensions will be assessed: throughput, measured as items completed per fortnight from the existing workflow system; quality, measured as the proportion of items returned for rework by the reviewing officer; cycle time from receipt to completion; and team experience, measured by a four-question survey run at the start, at three months and at the end. [MANAGER_NAME] pulls the first three. The survey is run by someone outside the team.
The baseline is the eight weeks immediately before the tool arrives, on all four dimensions. The stopping rule is that the pilot ends early if rework rises for two consecutive months, and expands if quality and cycle time both improve without a fall in team experience. Throughput staying flat is explicitly not a stopping condition.
That last sentence is the whole point. It costs nothing to write in advance. It is nearly impossible to introduce halfway through, once a flat throughput chart is on the screen and somebody senior is asking what the return was.
Bottom line
The evidence that AI helped your team may not appear in the number your tool exports, and it may not appear inside the window your steering committee allowed. A thirteen-month study of real delivery work found performance and perceived efficiency rising while activity stayed flat, which is the exact combination that a short activity-based pilot reports as no effect. Decide the window, the dimensions, the instruments, the baseline and the stopping rule before the tool arrives. After that, the data decides for you, and it will decide badly.
Do this Monday
- For every AI pilot currently running, write down its evaluation window and the dimensions it will be judged on, and note how many of those were decided before it started
- Add at least one non-count dimension to each, with a named instrument and a named owner
- Capture a baseline for anything not yet started, and accept that a pilot without one produces an opinion rather than a result
- Write the stopping rule and the expansion rule now, while nobody knows the answer
- State explicitly, in the pilot's terms of reference, that a flat activity measure is not by itself a reason to stop
Primary sources
- Tomaz, R., Guenes, P., Araújo, A. A., Baldassarre, M. T. and Kalinowski, M., Impacts of Generative AI on Agile Teams' Productivity: A Multi-Case Longitudinal Study, Proceedings of the 2026 IEEE/ACM Third International Conference on AI Foundation Models and Software Engineering (FORGE '26). https://doi.org/10.1145/3793655.3793728
- Tomaz, R. et al., Impacts of Generative AI on Agile Teams' Productivity: A Multi-Case Longitudinal Study, preprint, submitted 14 February 2026. https://arxiv.org/abs/2602.13766
TheAICommand. Intelligence, At Your Command.


