IA 360
Current Affairs

AgentFAIR uses AI agents to audit geospatial data

A new preprint describes 13 AI agents and a critic that assess whether geospatial datasets are findable, accessible, interoperable and reusable.

4 min read AI-generated Leer en español
AgentFAIR uses AI agents to audit geospatial data

On July 17, 2026, researchers Ming Chen and Pranav Pai introduced AgentFAIR, an AI-agent system meant to assess whether a geospatial dataset follows the FAIR principles: that it is findable, accessible, interoperable and reusable. The work is a preprint, so its findings are an early evaluation rather than a final certification of data quality.

FAIR is not algorithmic fairness

In this context, FAIR does not mean bias or fairness in an algorithm. It stands for Findable, Accessible, Interoperable and Reusable: qualities that help scientific data be discovered, understood by machines, combined with other sources and used again. Geospatial data make that task harder. A title and a licence are not enough; coordinate reference systems, spatial and temporal coverage, file formats and access services can all determine whether a dataset is actually usable.

The authors start from a practical problem. Existing automated checkers often disagree because they use different rules, evidence sources and scoring schemes. In a diagnostic study of 50 datasets from ten repositories, the standard deviation across normalized tool scores averaged 15 percentage points and reached 30.3 points for one dataset. That does not establish which tool is most accurate. It shows why a single score can conceal incompatible criteria.

Thirteen reviewers, plus a critic

AgentFAIR combines conventional metadata extraction with language models. It first loads a dataset landing page, including pages where some information is rendered with JavaScript, and gathers signals from HTML, JSON-LD, RDF/XML, DataCite and other formats. It then divides the assessment among 13 specialised agents, one for each FAIR sub-principle. Each agent gives a maturity score from 0 to 3, points to the evidence behind it and suggests an improvement.

Its distinguishing component is a critic agent. It does not judge a dataset on intuition: it checks whether evidence supports the score and whether separate conclusions conflict. When evidence is missing or confidence is low, it can request a targeted re-evaluation. The system produces a readable report, a JSON assessment, a SQLite evidence store and local traces, so a recommendation can be inspected rather than accepted as a black box.

What the study found — and what it does not prove

Across the sample, the paper reports average scores of 79.7% for findability, 70.4% for accessibility, 45.3% for interoperability and 72.0% for reusability. Interoperability was the weakest area: publishing a file is not enough if its metadata and conventions cannot communicate with other systems. In a repeatedly tested subset of ten datasets, sub-principle agreement was 89% with the critic and 71% when it was removed. The authors also report an API cost of about $0.054 per dataset.

Those are encouraging findings, but the paper states important limits. It covers only 50 datasets, comparisons with baseline tools are not a measure of accuracy, and validation used a single model family. Its expert review covered 15 datasets and reported 82% alignment, a useful signal but still a small basis for claims of broad generalisation.

The useful part is being able to discuss the score

The most interesting contribution is not another AI-generated overall grade. It is the prospect of turning a tedious audit into a checkable conversation: a licence is not expressed in a reusable form, a persistent identifier is missing, or a coordinate reference system is absent from a recognised standard. For people maintaining environmental, urban or climate datasets, that level of detail makes it possible to fix the specific issue.

The code is published under the AGPL-3.0 licence, and the repository notes that it does not include raw benchmark results or a comprehensive test suite. Its local service is also not designed to be exposed directly to the public internet. AgentFAIR does not replace expert review. Its more modest proposition is that agents can surface the evidence an expert needs to decide whether data are genuinely ready for reuse.

FAIR is not a universal quality grade

The original FAIR principles describe conditions that help data and metadata be found and reused by people and machines. They do not say that the content is correct, representative or suitable for every study. A dataset may have a persistent identifier, licence and interoperable format while containing biased measurements. Another may be scientifically valuable yet score poorly because its documentation is inadequate.

The four dimensions should therefore be read separately. Findable asks whether identifiers and metadata allow discovery. Accessible asks through which protocol and under what conditions the resource can be retrieved. Interoperable examines vocabularies, formats and references that other systems can interpret. Reusable requires provenance, licence, context and enough detail for another use. Offsetting one serious weakness with a high average would conceal the work that an audit is meant to surface.

How to audit the auditor

The first test is evidence preservation. Every score should point to the field, document or service response that produced it, with the access date. If a page changes, the team needs to know what the system saw. The AgentFAIR preprint makes traceability part of the design, but it only helps when citations are specific and dynamic material can be inspected again.

The second test is fixing the rubric before viewing the result. Two tools may use the same word while requiring different proof. One may accept a licence written in free text; another may demand a standard identifier. Disagreement is not resolved by choosing the higher grade. It is resolved by publishing the evidence required for each level and testing known cases that should pass or fail.

The third is separating consistency from accuracy. Repeating an evaluation and receiving similar answers shows stability, not correctness. Comparison with experts helps, but it also needs a shared guide and a record of their disagreements. The archived software on Zenodo fixes one artifact version; future evaluations should also fix models, prompts and dependencies so an improvement or regression can be attributed.

The fourth is testing missing and adversarial inputs. A page may contain contradictory metadata, a visible but machine-unreadable licence, broken links or malicious instructions aimed at the agent. An evaluator navigating external content must treat it as data, not instructions. The paper includes a local security boundary; exposing the service publicly would require additional isolation, limits and review.

The transferable skill is to reject an overall score until its evidence can be followed. For any automated data auditor, ask for the rubric, version, traces, test cases and a separation between consistency and accuracy. Good automation does not end the argument over a grade; it makes that argument reproducible.

The most useful assessment result may be “we do not know.” When a licence cannot be found, the system should distinguish an absent licence, an inaccessible page and a parser unable to read it. Those states lead to different actions and should not receive the same confident explanation. Properly recorded uncertainty helps prioritise human review. Invented certainty turns a limitation of the evaluator into an accusation against the dataset. A good report therefore includes not only scores and recommendations, but also collection failures, unresolved conflicts and the evidence the agent expected but could not obtain.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close