The Summary Kept the Number and Lost the Caveat, practitioner guidance from TheAICommand
← Leadership
Decision-making

The Summary Kept the Number and Lost the Caveat

Almost everything now reaching a senior leader has passed through a model that compressed it. New research on financial source material names the failure mode precisely: not fabrication, but decontextualisation, where the salient number survives and the qualifier that made it interpretable does not. The output reads well, agrees with the source, and quietly changes the decision. The fix is not to ban the summary. It is to change what you ask for.

Leading with AI. Written for Australian managers and people leaders. General information only. The judgement stays yours.

Quick answer

Compression is now the default path information takes to a leader, and the documented failure is decontextualisation: evidence retained, caveats stripped, output fluent. Ask for the source alongside the summary, require a line naming what was left out, run a second model over anything consequential and read the disagreements, and keep one channel that nothing summarises.

Your summary kept the number and dropped the caveat.

Somewhere between the source document and your inbox, most of what reaches a senior leader now passes through a model. The board paper draws on a compressed version of the operational report. The operational report draws on a compressed version of the incident log. The pre-read for tomorrow was written from a transcript nobody watched. Each step is defensible on its own, and none of them was a decision anyone announced.

The risk in that chain is not the one leaders have been trained to look for. It is not that the model made something up. It is that everything it kept is true, and the part that made the true thing interpretable is gone.

What the research actually found

A preprint published on 28 June 2026 and revised on 8 July 2026 puts a name to it. In "When Summaries Distort Decisions: Information Fidelity in LLM-Compressed Financial Analysis" (arXiv:2606.29251), Lee and colleagues test what happens when large language models compress financial filings and earnings-call transcripts, and they frame the problem in terms a leader can use: information fidelity, where "compression loses fidelity when it changes the decision induced by the source".

Their finding is that "LLM-based compression can produce fluent and factually plausible compressed contexts that nevertheless alter downstream decisions". They isolate two diagnostic patterns.

The first is decontextualisation, described as the case "where salient evidence is retained but separated from the caveats and contextual qualifiers needed for correct interpretation". The number survives. The sentence that told you the number only holds under one scenario does not.

The second is model dependency, "where different compressors expose different views of the same source". Two capable models, given the same document, do not fail identically. They surface different things, which means the version you received is partly an artefact of which tool your analyst happened to open.

The authors also make a point that matters more each quarter: "In agentic systems, such losses may recur across intermediate steps and amplify throughout the decision process." Compression is rarely a single step any more. It is a chain, and each link compresses the last link's output rather than the original.

Two honest caveats before anyone quotes this in a steering committee. It is a preprint and has not been through peer review. It was run on financial source material, not on management reporting, so applying it to your board pack is a reasonable inference rather than a demonstrated result. That said, the mechanism it describes is not finance-specific, and the caveat you just read is exactly the kind of sentence a summary drops.

A frame split in two by a thin gold rule, the left half a dense page of text under warm light with one line glowing, the right half a single clean glowing line alone in the dark, gold and sky on deep navy
Decontextualisation. The evidence survives the compression. The qualifier that made it interpretable does not.

Why this is harder to catch than a hallucination

A fabricated fact is a solved problem in principle. Check the claim against the source and it fails the check. Verification tooling, review habits and every AI policy written in the last two years are built around exactly that motion.

Decontextualisation defeats all of it. Every sentence in the summary is in the source. The check passes. The distortion is in what is absent, and there is no line to point at. You cannot fact-check a silence.

It also does not feel like risk. The compressed version is shorter, cleaner and more confident than the original, because hedging is the first thing compression discards. It reads like better writing. In practice it is a document that has quietly converted a probability into a fact, and confidence in a paper travels straight into confidence in a room.

This is a different problem from the one we covered in make AI disagree with you before you decide. That piece is about a model that agrees too readily inside a conversation you are having. This is about the material arriving before the conversation starts, already shaped, by a process nobody in the room ran.

The older problem this accelerates

Organisations were filtering information upward long before any of this. The reluctance to pass on unwelcome news is one of the most durable findings in social psychology, described by Rosen and Tesser in Sociometry in 1970 as the MUM effect: people transmit good news readily and bad news slowly, or not at all.

Compression does not create that tendency. Our reading, and this is inference rather than a finding of the research, is that it lowers the cost of acting on it. Softening a report used to require a decision by a person who then had to live with the wording. Now the flattening happens by default, at no cost, and with no author to hold to it. Nobody chose to remove the dissenting view. The dissenting view was one paragraph in a thirty-page document and it did not make the cut.

That is the practical difference between a leader who is misled and a leader who is under-informed. Nobody in this chain is behaving badly. That is precisely why it will not correct itself.

Where it shows up first

Three places in a normal month, all of them consequential, all of them already compressed by the time you see them.

The incident summary. The hardest thing to preserve through compression is uncertainty, and the first hours of an incident are almost entirely uncertainty. A status note that reads "root cause identified as a configuration change" is a compressed version of "the leading hypothesis is a configuration change, two other explanations remain open". One of those supports a decision to stand down the response. The other does not.

The vendor or investment recommendation. Compression rewards the comparable and discards the conditional. Pricing, capability and timeline survive. The clause about what happens on exit, the assumption behind the savings figure and the reviewer who dissented tend not to.

The people paper. Engagement themes, performance calibration inputs, restructure options. This is where the MUM effect and compression compound, because the uncomfortable material is both the most likely to be softened by a person and the least likely to survive a model. It is also the category where the cost of being under-informed lands on someone other than you.

If you only change one thing, change it here first.

Four things to change about what you ask for

None of this requires reading more. It requires asking for different things.

  1. Ask for the source with the summary, and name it. Every paper that supports a consequential decision carries a link to what was compressed and one line saying what it was. You will rarely open it. The requirement changes the author's behaviour, because a summary written next to a citable source is written more carefully than one written into a vacuum.
  2. Require a what-was-left-out line. Two or three sentences: the caveats, the dissent, the range around the central number. This is the highest-yield item on the list. It cannot be written without going back to the source, and it converts an invisible omission into a visible, signed choice.
  3. Run a second compressor on anything consequential. The research authors propose generating multiple candidate compressions and auditing their disagreements against the original source. You can do the cheap version of that this week: put the same document through a different model, and read only where the two summaries disagree. Model dependency stops being a hidden defect and becomes a search tool for the paragraph that matters.
  4. Keep one channel nothing summarises. A standing skip-level, a monthly hour on the floor, four raw customer calls a quarter. Not for the content, which will be unrepresentative, but because it is the only input in your week that has not been optimised for your attention. Ask questions a summary structurally cannot answer: who disagreed, what did we stop doing, what surprised you.
A left to right flow of four gold pill nodes on deep navy reading source, omissions, second reading and raw channel, joined by one continuous flowing line
Four asks. None of them require the leader to read more, and all of them change what the author does.

What this looks like in practice

[TEAM] brings a proposal to expand [PROJECT]. The paper says the pilot delivered a 22 per cent improvement in cycle time and recommends scaling to three more sites.

Under the old ask, you interrogate the number. Is 22 per cent right, how was it measured, over what period. All good questions, all of which the paper can answer, because the number is the part that survived.

Under the new ask, the what-was-left-out line reads: "The 22 per cent covers the six weeks after go-live; weeks seven to ten were not included because two of the four teams paused. The operations lead does not support scaling until the pause is understood."

Nothing in the first version was false. The second version is a different decision. It took the author twenty minutes and it took you thirty seconds.

If the line comes back empty, that is information too. Either the source genuinely had no qualifications, which is rare, or the author has not been back to it.

A single lit desk in a wide dark room, one open page glowing under a lamp with the far walls receding into depth and one distant door standing open, gold on deep navy
The lit page is not the room. Keep one channel open that nothing has compressed for you.

What you cannot delegate

The judgement that a decision is consequential enough to deserve the source, the second reading and the argument is yours. No model will flag it, because the model cannot see the stakes, only the text. That is the same boundary we drew in you are no longer the smartest person in the room: the information advantage is gone, the framing and the accountability are not.

There is a cost, and pretending otherwise is how good practices die in month two. The second reading takes analyst time. The what-was-left-out line takes author time. Applied to everything, it recreates the workload compression was meant to relieve, which is the same trap we described in the review tax. Apply it to decisions that are expensive to reverse, and accept a clean summary for everything else.

For the mechanics of what a model is given to work with in the first place, context engineering covers the layer beneath this one.

Bottom line

The information reaching you has been compressed, usually more than once, by a process nobody decided to introduce. The documented failure is not invention but decontextualisation: the evidence survives, the qualifier does not, and the result reads better than the original. Ask for the source, ask what was left out, read where two models disagree on anything expensive, and keep one channel that has not been summarised for you. Then spend the attention you saved on the decisions that cannot be reversed.

Do this Monday:

  • Add a what-was-left-out line to the template for any paper going to your leadership meeting
  • Pick the most consequential paper on this month's agenda and read the source it was built from
  • Run that same source through a second model and read only the disagreements
  • Put one unsummarised channel in your calendar as a recurring commitment, not an intention
  • Tell your team the caveats are the part you want, so the incentive stops running the other way

TheAICommand. Intelligence, At Your Command.

Frequently asked questions

What is the actual failure mode in AI summaries for decision makers?
A June 2026 preprint on compressing financial filings and earnings-call transcripts frames it as information fidelity, where compression loses fidelity when it changes the decision induced by the source. The authors name two diagnostic patterns: decontextualisation, where salient evidence is retained but separated from the caveats and contextual qualifiers needed for correct interpretation, and model dependency, where different compressors expose different views of the same source.
Is this the same problem as AI hallucination?
No, and that is why it is harder to catch. A hallucination introduces something that is not in the source, so a check against the source finds it. Decontextualisation keeps only what is in the source. Every sentence verifies. The distortion sits in what is absent, and absence does not trigger a fact check.
What is the single highest-value thing a leader can change?
Require a what-was-left-out line on any paper that supports a consequential decision: the caveats, the dissent and the range that did not make the summary. It costs the author two minutes, it is impossible to write without re-reading the source, and it converts an invisible omission into a visible choice someone has signed.
Does this mean leaders should read the raw material themselves?
Not generally. Reading everything is the problem compression was introduced to solve, and reverting relocates the cost rather than removing it. The practical position is selective: accept compression for the routine, demand the source and a second reading for the consequential, and keep at least one recurring channel that nothing summarises.
Decision MakingAI at WorkInformation QualityExecutive ReportingLeadership PracticeJudgement
← Back to Leadership