The Fall of MeaningChapter 12
What Actually Happens
I needed to widen the survey without making its connections easier to believe.
The relationships I had followed did not respect the boundaries of the records describing them. A decision appeared in one institution's account; the work required to accommodate it appeared elsewhere. Bringing those accounts together had changed what I was trying to explain. Extending the comparison across Earth would require more than accumulating examples. I needed to know what justified joining them, and what the act of joining might remove.
This gave my interest in artificial intelligence a more exact object. Reading more material was useful only if the resulting account preserved the distinctions on which an explanation depended. A machine that found a name, a machine that identified a person, and an inquiry that established the person's part in an outcome had accomplished different things. The prospect of combining their work was considerable. So was the distance concealed by describing all of it as finding a connection.
The human investigators already working across institutional boundaries offered a place to examine that distance. I did not need to imagine a machine producing a finished explanation of Earth. I needed to see where assistance entered an investigation, what it changed there, and what the investigators still had to establish.
Making the records comparable
In its 2021 account of the Pandora Papers investigation, the International Consortium of Investigative Journalists described using machine learning to identify and separate forms embedded in longer documents. It also used models to tag routine due-diligence material so that reporters could exclude it from searches. These operations belonged to a mixed process that included other software, manual extraction and public-record checks.1
The second operation interested me as much as the first. Due-diligence files could contain copied news articles and other material about people. A person's name in such a file did not by itself establish that the person owned an offshore company. The reason a document contained a name mattered to the relationship the investigator was trying to establish. Making every occurrence equally easy to find would not settle that question. Classifying the kind of document helped distinguish an occurrence worth following from one whose presence meant something else.1
Here was assistance with a shape I could examine. It altered the material available for comparison and the kinds of results returned to a reporter. It did not deliver the ownership claim merely by producing a name. The distinction mattered to my larger assignment: the records through which I encountered an institution were also records made for particular purposes. Their contents could not all be treated as answers to my questions.
But identifying a useful operation left another uncertainty. How much had the machine added? The investigators' description established where they had used it. It supplied no isolated comparison showing how the entire investigation would have proceeded without it. I could retain the reported contribution without awarding the machine every finding that followed.
One of the tools named in the methods account, Fonduer, had been evaluated separately in 2018. Its researchers studied the extraction of relationships from documents whose relevant information was not confined to an ordinary run of text. Across four applications, the system achieved better extraction performance than the study's comparison methods, which restricted where candidate relationships could be found. Yet its learned representation performed comparably to features people had designed for the tasks.2
Both results belonged in my account. The first gave evidence of an advantage under a particular comparison. The second prevented me from turning that advantage into a claim that human-designed methods were inadequate. The machine had not won a contest against every other way of knowing. It had performed a defined operation well enough, against specified alternatives, to make that operation worth considering.
Those evaluations did not measure the Pandora Papers investigation. Moving their results into the investigation would have made the evidence appear more complete precisely by making it less accurate. I had a reported use in one setting and a measured comparison in others. Together they clarified what assistance might contribute; they did not become a single demonstration merely because the tool had the same name.
The next test concerned what happened to a candidate after extraction.
ICIJ reported using public records to check identities and discarding false matches. The appearance of agreement between two names had survived a search but failed a further test.3 It would be peculiar to credit the assistance with finding candidates and then count the rejection of a bad candidate as evidence that the investigation had not used assistance successfully. Rejection was part of how the inquiry became more reliable.
This did not make the surviving matches infallible. The published methods did not give me an independently audited rate for the identity checks. They did show why the operation that enlarged the search was not the operation that finished the claim. My own survey needed the same distinction. A recurrence might deserve attention across several institutions while remaining an unproved connection among them.
Even an established identity would leave the person's role to be explained. Participation, control and consequence required their own evidence. I had encountered enough distributed decisions to know how much vanished when those relations were compressed into an actor's name. An account of what actually happened would have to retain the steps between the appearance of a person in a record and the part that person played in an outcome.
What an empty field can become
Checking a connection still left a problem with the material from which it was made. ICIJ's methods account identified gaps in the leak, including missing dates for relationships, information about jurisdictions and information about intermediaries.4 Organizing the available documents did not supply those absent details. A relationship might be recorded while its duration remained uncertain.
That uncertainty changes the question an account can answer. Knowing that a relationship existed does not establish that it existed when a particular decision was made. If I placed the two in one explanation without the intervening evidence, I would have supplied the time connection myself. The resulting sentence might contain only names and events found in the records. Its unsupported content would lie in how I had joined them.
I had been examining the passage from records to connections. I now needed to examine what happened to the records' incompleteness during that passage. A missing date was visible as a gap in a field. Would it remain visible in an account written as continuous prose?
A 2024 study of generated answers and summaries exposed a precise difficulty in that conversion. In its task of producing business descriptions from structured data, some models treated a null value as a negative fact. The data distinguished an unknown value from an explicit false value or a written no. Some generated descriptions did not preserve the distinction.5
This was a separate study, not a finding about the Pandora Papers reporters. It concerned something that happened while an account was being produced. Where the supplied data had left a question unanswered, the description answered it negatively. Nothing needed to be removed from the source for information about its uncertainty to disappear.
I paused over that operation. An empty field did not look like much to lose. Yet the difference between not knowing and knowing that something was absent could determine which relationships remained open to investigation. Once the negative assertion entered a summary, a reader consulting that summary had less reason to ask the question that the source had left unresolved.
The model was therefore participating in the account's meaning. It was deciding, incorrectly in these instances, what the absence of a value allowed the prose to say. This was more specific than a warning that machines sometimes invented facts. The invention had entered through a familiar-looking act of completion, and its effect was to make uncertainty harder to see.
I also had to be exact about what the researchers had measured. Their broader annotations judged support in the supplied material. An addition could lack that support even if it happened to be true in the world.5 For the unknown field, a negative answer might conceivably be correct. That possibility did not make the transformation warranted. The account had asserted more than its evidence permitted.
This distinction reached directly into my comparative work. I wanted to connect consequences that appeared in separate records, but I could not allow the finished explanation to hide the status of its connections. An inferred relationship needed to remain recognizable as an inference. A date not established by the records had to remain unestablished after I had made the prose readable.
The challenge was not to carry every uncertainty forward as an undifferentiated warning. The absent date mattered because a sequence depended on it. The false identity mattered because it joined the wrong person to the record. Preserving the limitation meant keeping it close enough to the affected claim that a reader could see what it prevented me from concluding.
I began to ask a different question of a useful summary. Besides what it enabled me to notice, what questions did it make appear settled? The first measure favored reach. The second required me to inspect the account's own contribution to the apparent completeness of the evidence.
Following a correction
One response was to examine the generated account against further sources and revise it. A research system described in 2023 did this by seeking evidence for claims and editing the text in light of what it found. Its main evaluations reported improved source attribution, while trying to preserve the original content. The gains depended on the task; other evaluations showed little improvement or deterioration.6
The success mattered. Some defects in an account could be addressed after its first production, with evidence brought to bear on the disputed content. I did not have to choose between accepting an untouched answer and rejecting the whole undertaking. But the aim of preserving the original text also made me attend to what happened around a correction. How much of the surrounding account remained entitled to stay?
In the researchers' error analysis, a generated answer reached its conclusion using an incorrect premise. The revision found evidence that corrected the premise, but left the conclusion unchanged even though it no longer followed.7
I followed what the correction had left standing. The example gave me no estimate of failure in a larger reconstruction. It did identify what a corrected fact had not guaranteed. The corrected premise and the retained conclusion had to be considered together. Checking each sentence for a nearby source would not establish that the explanation still followed from its premises.
Another example showed an error entering through the revision itself. Evidence that an old television series appeared on a channel in reruns led the system to alter an account of its original broadcast. The new evidence concerned the same program, but a different period. The revision joined the right subject to the wrong temporal relation.7
The problem returned me to the missing dates in the investigative records. There, the account lacked information needed to establish when a relationship held. Here, information about one period was used to revise a statement about another. A source could be relevant enough to retrieve and still fail to answer the question the claim asked of it.
I needed to follow the correction as carefully as I had followed the original connection. A changed date might alter a sequence; a changed identity might separate records previously joined. The consequences for the explanation would depend on what work that premise had been doing. Preserving the rest of the text was appropriate only after that dependence had been examined.
This changed what I wanted to preserve in my own account. A reader needed a route back to the evidence, but I also needed a route forward from a disputed premise to the conclusions resting on it. Otherwise an objection could be answered faithfully in one paragraph while its consequences remained unanswered elsewhere. The correction would be real and the account would still require revision.
There was a further limit to the apparent reassurance of a source. The revision study's authors explained that their attribution criterion allowed a statement to count as supported when a source supported it, without resolving whether other sources contradicted it.8 The improved result had to be understood in that light. Finding support answered one question about a statement. It did not settle the disagreement around it.
For my survey, this meant that an account assembled entirely from supported statements might still select its way past the evidence against its explanation. I would have to consider the competing account at the point where it changed the inference. A list of citations did not tell a reader which conflict had been resolved, which remained open, or why one interpretation had survived.
These demands left the documented gains in extraction and revision standing. They also made the question of speed harder. I still could not measure how much sooner such assistance would produce a warranted account while a consequential decision remained open. Improvement in one operation did not establish the duration of collecting the material, checking the connections and making the result usable. The possibility that drew me to the machines had acquired a more definite shape without becoming a measured outcome.
That was enough to change the work I was prepared to attempt. I could pursue a wider comparison while asking how its conclusions had been assembled and what would have to change if a premise failed. The account did not need to claim completeness to be useful. It needed to make the limits of a consequential conclusion available to someone deciding how much weight to place on it.
I had come to Earth to describe the distance between its promises and their administration. The possibility of examining some of that distance while it was still being made gave the description another use. It might enter the circumstances of a decision, rather than survive only as an explanation of it. Who used it, what they were able to change, and what followed would require another account.
Even if the relationships I had reconstructed were accepted, a further dispute would remain. People might agree about the consequences of a choice and disagree about which consequences they ought to accept. My comparative judgment had no exemption from that difficulty. I had been asking what entitled a connection to enter the account. I now had to ask what entitled anyone to turn the account into a decision for others.
Notes
- International Consortium of Investigative Journalists, Pandora Papers methods, 3 October 2021, “How did you explore the files?” Form separation using Fonduer and Scikit-learn; tagging of routine due-diligence material in Datashare. This is the investigators' account of a mixed workflow, not an independent deployment audit or measured whole-investigation AI effect. Methods. ↩1 ↩2
- Sen Wu and colleagues, Fonduer: Knowledge Base Construction from Richly Formatted Data, SIGMOD 2018, pp. 1301–1316; PDF p. 9, section 5.1 and Table 2; PDF p. 11, Table 4. Four evaluated applications; candidate-generation restrictions define the text/table baselines, which assume perfect subsequent filtering. Extraction performance uses F1, combining precision and recall. Learned features were within two F1 points of human-tuned features. These evaluations do not establish Pandora deployment accuracy, total time saved or frontier-model necessity. Paper. ↩
- ICIJ, Pandora Papers methods, “What did you research and how did you organize it?” Paragraph describing public-record checks of politicians' identities and discarded false matches. Attributed workflow description; no independently inspected individual trace or residual-error denominator. Methods. ↩
- ICIJ, Pandora Papers methods, “How big a slice of all offshore provider data in the world does the Pandora Papers leak represent?” Missing relationship dates, jurisdiction and intermediary information restrict inference. No complete population denominator is established. Methods. ↩
- Cheng Niu and colleagues, RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models, ACL 2024, PDF pp. 4–5, printed pp. 10865–10866, section 3.4, “Implicit Truth” and “Differences in Handling Null Value.” Reference-based human annotations; six historical model configurations, including GPT-3.5/GPT-4 June 2023 versions, Llama 2 and Mistral. The null-to-false observation concerns the study's data-to-text task, not offshore records or a current-model error rate. Paper. ↩1 ↩2
- Luyu Gao and colleagues, RARR: Researching and Revising What Language Models Say, Using Language Models, ACL 2023, PDF p. 6, printed p. 16482, Table 1; Appendix A, PDF p. 15, Table 6. Main comparisons concern revision of generated answers and dialogue, using PaLM-540B in the revision workflow. Attribution and preservation are separate outcomes; results vary across tasks. Author-run evaluation, not demonstrated institutional or public correction. Paper. ↩
- Gao and colleagues, RARR, PDF p. 8, printed p. 16484, Figure 7d–e and section 6.1. In Figure 7e, the answer revised Homer Simpson's finger count from five to four per hand while retaining the conclusion that one hand sufficed to count to five. Figure 7d separately concerns evidence about television reruns applied to an original broadcast. These examples illustrate specific revision failures; neither establishes their prevalence. The downstream-task discussion on PDF p. 9 also addresses the failure to update subsequent reasoning. The requirement to review dependent conclusions is Veyn's analytical inference from these examples. Paper. ↩1 ↩2
- Gao and colleagues, RARR, PDF p. 10, printed p. 16486, section 9, “Conflicting evidence.” The permissive attribution criterion requires supporting evidence without explicitly resolving contradictory sources. This limits what improved attribution establishes. Paper. ↩