IA 360
AI Fundamentals

A Preprint Proposes Auditing the Trust Chain Between Humans and AI Models

A paper on arXiv argues that when people and language models shape a public claim, the risk is not just a false answer: sources, inferences and responsibility must remain reconstructable.

5 min read AI-generated Leer en español
A Preprint Proposes Auditing the Trust Chain Between Humans and AI Models

On July 8, 2026, Mihnea C. Moldoveanu and Joel A.C. Baum posted a 50-page preprint to arXiv with an unfriendly title but an everyday question: how do we know what deserves trust when a public claim passes through people, institutions and language models? Their answer is not a lie detector or a new AI tool. It is a conceptual framework for examining the chain that supports a claim. The paper — classified under artificial intelligence and under social and information networks — is theoretical, not experimental: its authors "sketch" a language and mechanisms, they do not measure a reduction in errors.

The paper, Adversarial Social Epistemology for Assemblies of Humans and Large Language Models, starts from a familiar fact. A seemingly simple answer can contain an original source, a human summary, an institutional rule, an inference and a model-generated reformulation. Each step can introduce an error, omit a limit or present a conclusion as a fact. When the chain is hidden, the confident tone of the last sentence can improperly replace evidence. The paper names the four scaffolds a public claim rests on: testimony (someone said it), inference (from A, B follows), institutional certification (an authority backs it), and tacit trust (we take the rest for granted). An attack on trust does not need to break all four: it is enough to undermine one without anyone noticing.

The transferable skill is simple: do not begin by asking “did an AI or a person say this?” Ask “what chain produced this sentence, and can I inspect it?” That habit remains useful when the model, platform and topic change.

A claim does not arrive alone

The authors call their proposal adversarial social epistemology. “Social” means public knowledge depends on other people: witnesses, experts, media, databases, certifiers and readers. “Adversarial” does not mean everyone lies. It means that in settings with reputational, economic or political incentives, someone may have room to exaggerate, select data, omit conditions or attribute a conclusion to a source that does not support it. The machinery the authors borrow from philosophy of language sharpens the idea: every assertion carries a commitment — the duty to defend it if challenged — and an entitlement — others' right to rely on it. The characteristic abuse is to collect the entitlement without paying the commitment: to present something as backed so that others trust it, while dodging the obligation to sustain it when someone asks. A language model, which produces sentences with poise but with no one behind them to answer, is a perfect vehicle for that mismatch.

The paper argues that familiar labels —bubble, echo chamber or misinformation diffusion— are not always enough. The problem can sit inside a chain that looks legitimate: a real quote loses context; a reliable report is summarised as if it guaranteed a decision; an AI system combines documents without preserving which part supports which sentence. Trust fails not only because content is false, but because auditability is lost. There, the paper argues, is the real target of the attack: not the truth of a sentence, but the ability to trace it. Its technical proposal — epistemic networks enriched with an inferentialist semantics — is, translated, a way to draw who relies on whom and to require each link to keep what its claim follows from. It sounds abstract; its practical consequence is not: it forces a chain to declare its joints instead of presenting itself as a smooth piece.

That matters for language models because they make fluent prose easy to generate. A model can accurately summarise a source while failing to say whether its answer also contains its own inference, a generalisation or material it cannot check. A user may copy that text into an email, report or post. At the end of the chain, no one knows who asserted what or under which limit.

Why "echo chamber" is not enough

The contrast with the usual labels is what makes the framework useful. An echo chamber explains why we keep seeing the same kind of message; a bubble, why the opposite never reaches us; misinformation, why something false circulates on purpose. None of the three describes the case the authors worry about: a chain where each link is, on its own, defensible — the source exists, the summary is faithful, the institution is real — and yet the result misleads, because somewhere it was lost which part backed which. The failure is not in a lying node; it is in the joints between honest nodes.

That shift of focus matters with language models precisely because they are machines for smoothing joints. A good model stitches a primary source, an intermediate report, and a conclusion into a terse paragraph where the seam no longer shows — and that smoothness, which we read as quality, is exactly what erases the information the framework wants to keep. Fluency is not neutral: it is the solvent of auditability.

Five questions that open the chain

First, source: where does this specific fact come from? A link to a general page is not enough when a claim depends on a table, date or exact sentence. Second, transformation: was it copied, summarised, translated or combined with other sources? Every transformation can change scope.

Third, inference: what is written in the source, and what is a conclusion? “The company filed a document” and “the company will do X” are not the same claim. Fourth, responsibility: who chose the source, wrote the summary, approved model use or decided to publish? Naming responsibility is not an automatic search for blame; it makes correction and explanation possible.

Fifth, review: can another person retrace the route and reach a reasonable answer from the same material? If there is no document, version, quotation, record or correction path, the claim deserves proportionally less confidence. This checklist does not prove a sentence false. It makes visible what would need checking.

AI as assistant, not certificate

A model can help at several points: locate documents, compare versions, extract quotes or draft a synthesis. None turns its output into a certificate. A person using the result should keep the links, distinguish quotation from interpretation and explain why a source fit the question.

There is a symmetrical trap too: asking the model itself to “verify” its answer does not replace independent review. It may check format or compare accessible sources, but readers cannot assess the result if the material and method remain hidden. Useful verification leaves reviewable traces.

A proposal, not a demonstrated solution

The paper should not be sold as more than it offers. It is a preprint, submitted to arXiv on July 8, not an experimental study measuring error reduction or a deployed system. The authors outline language and mechanisms for auditing trust breaches. The idea can inform product design, editorial processes and future research; it still needs applications and evaluation.

That limit does not make it irrelevant. Many AI-supported decisions fail before technical evaluation: an answer is accepted because it reads well, cites an institution or cannot be reconstructed in time. Making the chain part of the product —with sources, changes, responsibilities and appeal paths— is a design choice available now.

The next time a system gives you an important conclusion, do not choose between total trust and total rejection. Ask for the chain. If you can distinguish source, transformation, inference, responsibility and review, you have something better than a reliability promise: a way to test it.

Primary source: arXiv preprint and PDF.

Sources for this piece

This piece draws on 3 primary source(s), gathered during reporting.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close