IA 360
Current Affairs

UK government tests AI for routine planning applications

Barnet, Camden and Dorset are testing an assistant that organises case files and drafts initial assessments. Decisions remain human; the test is traceability, accuracy and real time savings.

4 min read AI-generated Leer en español
UK government tests AI for routine planning applications

The promise sounds simple: reduce the wait for a household extension from eight weeks to four. The test is much harder. On 16 June 2026, the UK government introduced an artificial-intelligence prototype to assist with routine planning applications in Barnet, Camden and Dorset. The system can assemble documents, locate rules and prepare an initial assessment; a qualified officer retains the decision. Whether that boundary works teaches a durable lesson about AI in government: automating steps is not the same as delegating authority.

A pilot, not automatic planning permission

The official announcement calls the project Augmented Planning Decisions, or APD. The prototype triages applications, summarises relevant information and produces a preliminary assessment for the officer. Initial testing focuses on householder cases such as extensions, loft conversions and conservatories. The government did not announce an algorithm that grants permission by itself, nor a national deployment that had already been decided.

The government’s target is to cut the average processing time for a straightforward case from eight to four weeks. At the date of this article, that figure was an ambition, not a measured result. The Housing Ministry’s Digital Planning team said the alpha phase had begun in May and that the prototype could change significantly. Only if it showed value by the end of the summer would officials decide whether to extend testing to nine more local planning authorities for twelve months. Any nationwide expansion in 2027 would also depend on evidence.

The geography deserves precision. Barnet and Camden are London boroughs; Dorset is a local authority. The three were selected to expose the tool to different planning environments and systems, not because it had already passed a representative test across England. The pilot can reveal whether one product adapts to three workflows; it cannot establish in advance that performance will be identical in every council.

What the machine does and what the person must retain

The fuller description separates several tasks. The assistant pre-processes files, highlights missing data, extracts site information, finds national and local policies, reviews consultation comments and drafts an assessment report with reasons and possible conditions. Google DeepMind, a technology partner, adds that the officer reviews every line, changes the reasoning and retains the power to approve or reject.

That division shows where risk appears. Finding a policy does not prove it applies; summarising an objection can omit its decisive condition; proposing a condition does not prove it is necessary or lawful. Even literal extraction can attach a plan, parcel or date to the wrong case. Human oversight therefore cannot mean clicking “accept”. The reviewer needs the original document, cited provision, version used and route connecting each fact to the recommendation.

DeepMind says the prototype records its work at every step and creates an audit trail. That description is a property to test, not a sufficient guarantee. A useful trail need not expose a model’s unknowable internal reasoning; it should preserve inputs, retrieved documents, transformations, outputs, officer edits and the system version. Those elements make an error reproducible and keep the automated proposal distinct from the administrative decision.

The ministry draws another boundary: balancing competing considerations, understanding local context and weighing effects on homes and communities remain with the officer. “Human in the loop” protects the public only if that person has the time, competence and authority to disagree. If the projected saving depends on rapidly approving a plausible draft, automation may move the bottleneck from writing to verification.

Do not confuse APD with Extract

The announcement presented a second product, Extract, at the same time, and merging the two exaggerates maturity. Extract turns historic planning documents and maps—including files with handwritten notes—into reusable data. APD uses case information to support application processing. One prepares the documentary layer; the other participates in case analysis. Extract’s availability to English councils did not mean APD was available to them.

After trials with twenty local authorities, the government estimated that Extract could save an average council about 255 hours of manual work a year, down from more than 500. The announcement combines an observation from trials with an expected saving and does not provide the full measurement design. It cannot be transferred automatically to APD. Each product needs its own baseline, case sample, error rate and total review time.

The distinction also helps diagnose failures. If a rule was never digitised, the input data is at fault. If the correct document exists but search fails to retrieve it, retrieval is at fault. If it is retrieved but the summary changes its meaning, generation is at fault. If the recommendation is correct but nobody can explain its basis, governance is at fault. Saying only that “the AI was wrong” hides four different problems with different remedies.

How to evaluate a public-sector assistant

The first metric is end-to-end time, not how quickly the model produces text. Measurement should run from receipt of a complete file to communication of a decision, including corrections and review. Counting only minutes saved on a first draft omits work that returns later. Cases must also be comparable: a simple, complete application is not a valid baseline for one with inconsistent plans or a heritage constraint.

The second metric is quality by task. What proportion of relevant policies does it retrieve? How many citations point to the correct provision? Which objections disappear in a summary? How often does the officer change proposed conditions? Averages can hide harm concentrated in neighbourhoods, housing types or uncommon cases. Recording errors by category shows whether time is saved by treating outliers worse.

The third metric is contestability. A resident must be able to know what evidence was considered, understand the reason for the decision and use ordinary review or appeal routes. The Planning Inspectorate’s guidance on AI in casework evidence requires disclosure when a tool has drafted, summarised or analysed material and makes the submitting party responsible for checking accuracy. Although that guidance addresses material submitted to Inspectorate appeals and cases, it expresses a transferable rule: the provenance of a summary matters when another person must judge its reliability.

The fourth metric is operational change. The ministry’s own guidance on digital planning tools warns that data quality and coverage constrain results, accuracy can vary by site type, and outputs should be validated before informing decisions. It also calls for configured criteria, documented governance and continued professional review. A strong model placed on incomplete records and unclear responsibilities does not create a strong service.

The evidence still missing

As of 25 July 2026, APD remained an alpha with three authorities. The open sources had not reported results demonstrating an eight-to-four-week reduction, task-level accuracy, differences between groups or the full cost of verification. The ministry’s £8.2 million contract, with Google Cloud, Google DeepMind and Faculty as delivery partners, funds development; its price is not evidence of effectiveness.

That does not make the experiment pointless. Householder applications represent nearly 70% of the workload described by the government, so organising files and locating rules may release capacity if the system is accurate and oversight is real. But the meaningful result is not that the prototype can draft a convincing report. It is that cases move with less delay without increasing errors, concealing evidence or weakening the right to an explanation.

The transferable capability is to separate four layers in any administrative AI: data, technical assistance, professional judgement and legal authority. When an announcement says “a human decides”, ask what that person can see, how much time they have to check it, what is recorded and how an outcome can be corrected. That is the difference between oversight as a slogan and demonstrable accountability.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close