IA 360
AI Fundamentals

A Preprint Proposes Auditing the Trust Chain Between Humans and AI Models

A paper on arXiv argues that when people and language models shape a public claim, the risk is not just a false answer: sources, inferences and responsibility must remain reconstructable.

5 min read AI-generated Leer en español
A Preprint Proposes Auditing the Trust Chain Between Humans and AI Models

On July 8, 2026, Mihnea C. Moldoveanu and Joel A.C. Baum posted a 50-page preprint to arXiv with an unfriendly title but an everyday question: how do we know what deserves trust when a public claim passes through people, institutions and language models? Their answer is not a lie detector or a new AI tool. It is a conceptual framework for examining the chain that supports a claim.

The paper, Adversarial Social Epistemology for Assemblies of Humans and Large Language Models, starts from a familiar fact. A seemingly simple answer can contain an original source, a human summary, an institutional rule, an inference and a model-generated reformulation. Each step can introduce an error, omit a limit or present a conclusion as a fact. When the chain is hidden, the confident tone of the last sentence can improperly replace evidence.

The transferable skill is simple: do not begin by asking “did an AI or a person say this?” Ask “what chain produced this sentence, and can I inspect it?” That habit remains useful when the model, platform and topic change.

A claim does not arrive alone

The authors call their proposal adversarial social epistemology. “Social” means public knowledge depends on other people: witnesses, experts, media, databases, certifiers and readers. “Adversarial” does not mean everyone lies. It means that in settings with reputational, economic or political incentives, someone may have room to exaggerate, select data, omit conditions or attribute a conclusion to a source that does not support it.

The paper argues that familiar labels —bubble, echo chamber or misinformation diffusion— are not always enough. The problem can sit inside a chain that looks legitimate: a real quote loses context; a reliable report is summarised as if it guaranteed a decision; an AI system combines documents without preserving which part supports which sentence. Trust fails not only because content is false, but because auditability is lost.

That matters for language models because they make fluent prose easy to generate. A model can accurately summarise a source while failing to say whether its answer also contains its own inference, a generalisation or material it cannot check. A user may copy that text into an email, report or post. At the end of the chain, no one knows who asserted what or under which limit.

Five questions that open the chain

First, source: where does this specific fact come from? A link to a general page is not enough when a claim depends on a table, date or exact sentence. Second, transformation: was it copied, summarised, translated or combined with other sources? Every transformation can change scope.

Third, inference: what is written in the source, and what is a conclusion? “The company filed a document” and “the company will do X” are not the same claim. Fourth, responsibility: who chose the source, wrote the summary, approved model use or decided to publish? Naming responsibility is not an automatic search for blame; it makes correction and explanation possible.

Fifth, review: can another person retrace the route and reach a reasonable answer from the same material? If there is no document, version, quotation, record or correction path, the claim deserves proportionally less confidence. This checklist does not prove a sentence false. It makes visible what would need checking.

AI as assistant, not certificate

A model can help at several points: locate documents, compare versions, extract quotes or draft a synthesis. None turns its output into a certificate. A person using the result should keep the links, distinguish quotation from interpretation and explain why a source fit the question.

There is a symmetrical trap too: asking the model itself to “verify” its answer does not replace independent review. It may check format or compare accessible sources, but readers cannot assess the result if the material and method remain hidden. Useful verification leaves reviewable traces.

A proposal, not a demonstrated solution

The paper should not be sold as more than it offers. It is a preprint, submitted to arXiv on July 8, not an experimental study measuring error reduction or a deployed system. The authors outline language and mechanisms for auditing trust breaches. The idea can inform product design, editorial processes and future research; it still needs applications and evaluation.

That limit does not make it irrelevant. Many AI-supported decisions fail before technical evaluation: an answer is accepted because it reads well, cites an institution or cannot be reconstructed in time. Making the chain part of the product —with sources, changes, responsibilities and appeal paths— is a design choice available now.

The next time a system gives you an important conclusion, do not choose between total trust and total rejection. Ask for the chain. If you can distinguish source, transformation, inference, responsibility and review, you have something better than a reliability promise: a way to test it.

Primary source: arXiv preprint and PDF.

Sources for this piece

This piece draws on 3 primary source(s), gathered during reporting.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close