Torrejón tests AI to read safety patterns, not to watch people
Polic-IA turns historical incidents into risk maps. Audit data, thresholds, feedback, and human use before trusting a forecast.
On June 3, 2026, the city council of Torrejón de Ardoz presented Polic‑IA at South Summit Madrid, a web tool that estimates risk by area, date, and time slot from historical Local Police incidents. The distinction that must survive from the first line is this: the project says it analyzes patterns rather than identifying people, conducting automated surveillance, or deciding for an officer. A product name cannot prove that limit. It must be checked through the data, output, actual use, and ability to influence a decision.
The council’s official announcement, published on July 2, credits the project to UCAM Computer Engineering student Álvaro Juárez García as his final-degree project. Alberto Hernández, the municipality’s head of New Technologies, presented it. The application lets a user select a place and time, obtain an estimate, and visualize the historical distribution of incidents. The announcement does not publish the dataset, architecture, metrics, thresholds, or operating protocol.
A map of records is not a complete map of risk
The system starts from recorded incidents, not every incident that occurred. That small-looking distinction defines what a model learns. A police database reflects reports, calls, officer presence, classification criteria, and administrative changes. Two areas with the same underlying behavior may produce different records if one receives more patrols, residents report at different rates, or an incident is coded differently.
A historical concentration therefore supports a limited description: more cases were recorded there, under these practices, during this period. It does not support a direct leap to “people here are more dangerous” or “a crime will happen here.” Removing names is not enough either. Even when an output is an aggregated map, a territorial signal may affect people who live, work, or travel in each area differently.
The first transferable skill is to follow data provenance. For every point on the map, an evaluator should know which event counts, who records it, when it enters the system, how duplicates and category changes are handled, how much history is retained, and what remains missing. If those rules change, the time series is no longer directly comparable. A model does not repair a recording system by itself; it learns the version it receives.
An estimate needs a reference, horizon, and threshold
“Risk level” is not yet a reproducible variable. The target must be defined: a call, a report, or a validated incident, and within which horizon. Spatial units must also be specified. A very small grid may generate unstable values; a large area may conceal internal differences. A prediction for one hour and another for one week should not be judged against the same reference.
Then comes the threshold. A continuous score becomes categories such as low, medium, or high through selected cutoffs. Moving them changes how many areas are flagged and how many cases are missed. Performance should expose false positives and false negatives, calibration, and a comparison with simple baselines such as recent frequency. A more attractive map is not necessarily a better forecast.
Temporal validation is essential: train on the past and test on a later period the model could not know. A random split may place nearby or repeated events on both sides. Results should also be separated by area and period because a municipality-wide average can conceal systematic failure in one neighborhood or during unusual events. The municipal announcement gives none of these results, so the public evidence does not establish Polic‑IA’s accuracy or superiority over another planning method.
Use can create the data used for later training
A particular risk appears when a prediction changes where observation occurs. If a map concentrates patrols in selected areas, more incidents may be recorded there even if underlying behavior has not increased. Those records return to the system and reinforce its initial signal. Ensign and colleagues’ work on runaway feedback loops in predictive policing formalized how this process can amplify a biased distribution of observation.
This does not prove that Polic‑IA creates such an effect; its deployment is not public. It does provide a concrete test for any municipality. Indicators that depend on police presence should be separated from more independent measures, comparison areas should be preserved, changes in recording after resource changes should be examined, and model outputs should not automatically become new training truth.
Human involvement does not remove the issue by itself. A professional may receive a score without its uncertainty, interpret “high” as certainty, or lack time to challenge it. Meaningful human review means the person understands the scope, has alternative information, can override the suggestion, records why, and remains accountable for the decision. If a map determines resource allocation in practice, calling it “support” does not change its effect.
Legal classification depends on purpose and actual use
The European Union Artificial Intelligence Act includes in Annex III certain systems intended for law-enforcement authorities to assess the risk that a person becomes a victim or commits an offense. Article 6 allows some Annex III uses not to be treated as high-risk when they pose no significant risk and are limited, for example, to a preparatory task that does not replace or influence a human assessment without appropriate review. The exception must be documented and does not operate in the same way when a system profiles people.
On May 19, 2026, the Commission published draft classification guidelines. Their examples are not exhaustive and the document remained a draft. They help interpret intended purpose, preparatory roles, and material influence; they do not certify a particular product from outside. Public information is insufficient to assign Polic‑IA a definitive legal category.
Separate from system classification, data processing by competent authorities has its own framework. Directive (EU) 2016/680 covers personal data processed for the prevention, investigation, detection, or prosecution of criminal offenses. Aggregation or anonymization may reduce risk, but documentation should clarify which data enter, for what purpose, and who receives access. “Does not identify people” is a relevant claim that should be testable in the design rather than only stated during a presentation.
What a responsible municipal pilot should publish
A useful public record need not expose sensitive operational data. It can describe the objective, period, and origin of records; spatial and temporal units; excluded variables; validation method; metrics with intervals; limits of use; authorized users; logs of queries and overrides; review frequency; and the procedure for suspending the system. It can also say whether the pilot will alter resource distribution and how that effect will be measured.
Evaluation should separate three outcomes. The technical result asks whether an estimate matches the defined reference. The operational result asks whether it improves planning information or reduces workload. The public result asks whether safety improves without shifting surveillance or friction without justification. Better technical performance does not guarantee the other two. Publishing failures and unevaluable cases would be as informative as publishing successes.
Polic‑IA arrives with a sensible verbal boundary: historical patterns, no identification, and no autonomous decision. The next test is turning that boundary into observable controls. Readers can apply five questions to any similar announcement: what was recorded, what is predicted, which action changes, how error is measured, and what prevents a feedback loop. With those answers, a map stops looking like an authority and returns to what it should be: a partial representation that someone must justify.
Sources for this piece
This piece draws on 3 primary source(s), gathered during reporting.
This article was produced with artificial intelligence under human editorial oversight.