In the years before the 2003 invasion of Iraq, American intelligence agencies accumulated scores of reports about mobile biological weapons laboratories.
Much of that reporting came from one Iraqi defector, codenamed Curveball.
Between January 2000 and September 2001, the Defense Intelligence Agency disseminated almost 100 reports based on his accounts. Yet the 2002 National Intelligence Estimate attributed its judgment about mobile facilities to “multiple sensitive sources.”
The Commission's report exposed the imbalance behind that description. Two other sources had contributed one report each. Another source had already been judged a fabricator. The assessment depended overwhelmingly on Curveball, whose own reporting was ultimately determined to be fabricated. Presenting the judgment as supported by multiple sources obscured how narrow its foundation actually was.
The reports multiplied. Independent corroboration remained thin.
Most of us will never work on decisions carrying consequences remotely comparable to war. But the information failure underneath this story is surprisingly ordinary.
Someone says something. Someone else records it. Another person summarizes it, and the summary enters a product brief. The brief informs documentation. The documentation feeds customer support. An AI system retrieves that documentation and answers a user’s question instantly, confidently, and with citations.
At that point, how many sources do we have?
And how many things do we actually know?
When the Meaning Changes
Imagine an engineering team documenting the behavior of a data-transfer feature.
Engineering records a narrow, conditional behavior: “Scheduled transfers retry up to three times after temporary connection failures.”
Product simplifies it to: “Failed transfers retry automatically.”
By the time it becomes customer-facing documentation, the claim is broader: “No intervention is required when a transfer fails.”
An AI system can then repeat that final statement while citing several authoritative-looking sources.
The engineer describes the behavior precisely. The product manager removes implementation details because customers do not need to know everything engineering knows. The writer wants to answer the customer’s actual question without drowning the reader in qualifications.
Each person may be making a defensible editorial decision.
And somewhere during those decisions, the meaning changes.
“Scheduled,” “up to three times,” and “temporary connection failures” disappear. A limited retry behavior becomes an assurance that the customer does not need to intervene.
A better simplification would preserve those conditions: “Scheduled transfers automatically retry up to three times after temporary connection failures.” Anything broader requires broader evidence.
No one necessarily lied or misunderstood the product. The organization simply kept making the sentence easier to use.
The awkward part is that I agree with the instinct to simplify. I do not want a help article to become an inventory of everything Engineering knows. But I also do not want a shorter sentence to describe a more capable product.
Some details are clutter. Others are the conditions under which the statement is true.
The difficult editorial question is whether I made the same claim easier to understand or quietly made a different claim.
The problem compounds as that revised claim moves through the organization. Product knowledge gets compressed into shorter explanations, propagated across documentation and support materials, and promoted from tentative explanation to approved guidance.
Those are useful transformations. They are how organizations put knowledge to work. But they can also make a claim appear more authoritative without making it better supported.
Publishing the retry statement does not make the system retry more kinds of failures.
But We Have a Source of Truth
There is an obvious response to all of this:
Isn’t this what source-of-truth systems are supposed to solve?
Sometimes, yes.
Documentation platforms can reuse shared content across pages so teams do not have to maintain slightly different versions of the same explanation. We use these capabilities to keep shared explanations consistent and distribute corrections. But a reusable block can spread an overbroad statement just as efficiently as an accurate one. Keeping the copies aligned does not establish whether the claim is supported.
Consistency asks: Are we saying the same thing?
Support asks: What entitles us to say it?
Those questions are related, but they are not interchangeable.
Even formal provenance works this way. The W3C defines provenance as information about the entities, activities, and people involved in producing something, which can help us assess its reliability and trustworthiness. Provenance gives us something to evaluate. It does not magically turn the thing being traced into truth.
A beautifully traceable mistake is still a mistake.
There is another complication. Not every consequential statement is a hypothesis waiting for independent verification.
Some statements are authoritative because someone with the appropriate responsibility has made a decision. A policy becomes operative because the organization adopted it. A specification can define intended behavior. A product commitment can become organizationally binding because the organization approved and published it.
In those cases, asking for independent corroboration can be the wrong test. The question is whether the statement comes from the kind of authority capable of establishing it.
Claims about actual current behavior require a different test. If we say the system retries a transfer, revokes access, deletes data, or stops an automated action under particular conditions, an approved sentence does not establish that the system actually behaves that way.
Different claims earn authority in different ways. The discipline is knowing what kind of support entitles the organization to make a particular claim.
Now Give the Problem to AI
Ask an AI-assisted system to compare statements about retry behavior across engineering records, product documentation, and support conversations. Have it identify lost qualifications and contradictions, then show what supports each version. Ask which sources independently confirm the behavior and which appear to repeat another source.
A hypothetical response might look like this:
The engineering material supports automatic retries for scheduled transfers experiencing temporary connection failures. Later materials generalize this behavior to failed transfers broadly. I found no evidence in the reviewed sources establishing that all transfer failures recover without intervention.
But now the same scrutiny has to apply to the AI. Did it establish that those later statements descend from the engineering note through an explicit link, revision history, or citation, or did it infer lineage because the statements sound similar?
Likewise, “I found no supporting evidence” is not the same as “supporting evidence does not exist.” The relevant verification may have occurred outside the systems the AI could access. Two people may even have observed the same behavior independently and described it similarly.
AI can make organizational knowledge far easier to interrogate. It can also make inference look suspiciously like discovery.
A useful system should therefore expose what kind of answer it is giving: a recorded relationship, an observed behavior, an explicit decision, an inferred relationship, conflicting support, or unavailable evidence.
A Human in the Same Loop
Now suppose the organization goes one step further.
A support agent answers a customer:
No intervention is required when a transfer fails.
A human reviews the response before it is sent. The reviewer checks the product documentation.
The documentation says:
No intervention is required when a transfer fails.
Approved.
The reviewer confirms that the answer repeats the approved guidance. That can be a legitimate conformity check. It does not independently establish the product behavior that guidance describes.
A reviewer can still contribute useful judgment while reading the same source, including spotting an unsupported inference or a missing qualification. When the task is to check an answer against an approved policy, that source may be exactly what the reviewer needs.
If the agent, reviewer, support article, and product summary ultimately depend on the same unsupported assertion, adding more participants to the loop does not create corroboration. It creates more confidence around the same ancestor.
For a consequential decision, the reviewer may need access to something different: the relevant test, current behavior, engineering confirmation, operational telemetry, an approved policy decision, or another form of support appropriate to the claim.
Otherwise, the loop closes perfectly around an assumption.
Before the Sentence
Technical writers encounter this problem in an unusually important place: the point where internal knowledge becomes an external explanation.
We perform transformations constantly. We take what engineering knows and make it useful to a customer. We compare what a product manager intends with what the product currently does. We combine decisions scattered across Jira, Slack, demos, calls, requirements, and the product itself into something another person can safely rely upon.
Writing is the visible output, but the difficult work happens before the sentence.
We have to determine what the sources actually support, then write without expanding the claim beyond those limits. That responsibility brings us to one question:
What will someone do because we wrote this?
That question matters more as documentation becomes part of systems that generate answers, recommendations, and actions at scale. Documentation has long fed systems beyond the person reading a page. Natural-language documentation can now influence more directly what those systems produce and, increasingly, what they do.
A sentence written for explanation can become an input to execution.
That changes the stakes of seemingly ordinary editorial decisions.
Where to Stop
The obvious overreaction is to trace everything, attach provenance to every paragraph, record confidence levels for every sentence, and turn a help center into a forensic evidence repository.
No one wants to work in that system, nor would most users benefit from it. The appropriate response is proportional.
Most sentences are not consequential enough to investigate, but some claims change what people do. A statement may cause someone to leave a process unattended, grant access, delete data, spend money, configure security, disregard an alert, or allow an automated system to proceed without intervention.
Those claims deserve more attention.
For those claims, five questions help identify what needs checking:
What exactly are we asserting?
What supports that assertion?
Which conditions or qualifications make that support applicable?
Has any downstream version increased the claim’s scope or certainty?
What decision, human or automated, might this statement authorize?
Sometimes the answer will be simple. Engineering confirms the behavior, a current test supports it under the relevant conditions, and the wording accurately preserves those constraints. Publish it.
Sometimes the questions expose a gap. Narrow the claim to what is supported, have the responsible team verify the broader behavior or clarify the policy, or hold the claim until it is checked. If a reusable block changes, downstream material that inherited it may need correction. If an agent relies on that material, its answers may need to be tested again.
The writer does not need to reconstruct the complete ancestry of every statement before publishing it. In many cases, establishing what is supportable now matters more than discovering who first said it or where it first appeared. Keep the support connected to the claim, along with the conditions or changes that would require another check.
The objective is justified reliance, not perfect ancestry.
Before the Echo
The Curveball case illustrates how abundant reporting can conceal a narrow and unreliable source base. The risk in our own work is mistaking agreement across documents, support answers, and AI responses for independent confirmation.
A claim can travel farther without becoming better supported.
When an organization asks a person or an agent to rely on a claim, someone should be able to answer a more difficult question than Where did we write this?
Why are we entitled to say it?
A mature knowledge system does not merely know what to say.
It knows why it is allowed to say it.




