Berlin prepares a forum to put public-health AI under the microscope
The Robert Koch Institute will hold its third symposium on AI in public-health research in September, focusing on data, evaluation and governance.
Artificial intelligence in public health is not settled by a single software demonstration. It needs comparable data, teams able to interpret results and rules that clarify what may be done with them. That will be the ground covered by the third “Artificial Intelligence in Public Health Research” symposium, which the Robert Koch Institute (RKI) will hold in Berlin on 9 and 10 September 2026.
Bringing research and practice together
The event is organised by the RKI’s Centre for Artificial Intelligence in Public Health Research, known as ZKI-PH. The Institute’s event page says it receives financial support from Germany’s Federal Ministry of Health through the AI-DAVIS-PANDEMICS project. That matters because the discussion is not framed as computer science alone: it starts with public-health functions, where data, procedures and public trust matter as much as the model.
The previous edition helps explain the approach. The RKI says the second symposium, held in May 2025, brought together more than 180 participants and placed 15 talks in four sessions: AI-supported public-health decision-making, strategies for antimicrobial resistance, climate change and public health, and regulatory frameworks for AI and machine learning. The third meeting continues a series that began in 2023.
This is not a catalogue of applications that are ready to deploy. It is a programme of questions: how to spot disease patterns earlier without mistaking a signal for a conclusion; how to make information from different institutions usable together; and how to evaluate a tool before its outputs enter a public-health decision.
Data are not a technical footnote
The RKI’s account of the 2025 meeting highlights a repeated need: structured, standardised data. Without that foundation, comparisons across regions, population groups or time periods can produce misleading conclusions. Public-health records also reflect how measurement is done, who reaches services and which changes have occurred in a particular surveillance system.
AI can help organise very large bodies of information, identify regularities and support specialist work. But its outputs need context. A model may identify an association worth investigating; it cannot decide on its own whether that association has an epidemiological explanation, whether data are systematically missing, or whether a public intervention is proportionate.
The World Health Organization sets out that balance in its guidance on ethics and governance of AI for health. It recognises potential for health research, surveillance and outbreak response, while placing ethics and human rights at the centre of design, deployment and use. Its principles include preserving human autonomy, promoting safety and wellbeing, ensuring transparency, fostering accountability, supporting equity and evaluating systems over time.
What needs to be measured
That is why a meeting such as Berlin’s matters more for the questions it organises than for the promises it collects. Before a system enters a public routine, it is worth knowing which population it was evaluated on, the quality of its data, when it fails, who reviews its recommendations and how an outcome that harms or excludes someone can be corrected.
The symposium will bring those conversations together before AI is treated as an automatic answer. In public health, technology can extend the capacity to observe and analyse. Responsibility for interpreting, explaining and acting will remain collective and human.
How to read an evaluation before turning it into policy
The first check is to separate the public-health objective from the statistical objective. Predicting who may miss an appointment, detecting a change in consultations and estimating the spread of an infection are different tasks. Each needs an explicit outcome, time window and downstream decision. A model may perform well on its metric and still be useless if it delivers the signal too late, if no available action changes the outcome, or if acting on it diverts resources from a better intervention.
The second check is to find the denominator. “It detected more cases” does not say how many people were examined, what share truly had the condition or how many false alerts the team received. Sensitivity and specificity describe different aspects, but neither is sufficient alone: the value of an alert changes with prevalence and with the cost of reviewing it. An evaluation should show an outcome table that the service expected to act can understand, not just one aggregate number that favours the model.
The third check is time. Health data change when a diagnostic test, case definition, campaign or a population’s access to care changes. That drift can damage a tool without any change to its code. The RKI’s work on AI and public health therefore cannot stop after a single training run: performance monitoring, review dates and a rule for withdrawal or recalibration are part of the system.
The fourth is to ask who is missing from the data. A record may underrepresent people who do not reach a clinic, cannot use a digital service or appear with incomplete fields. When a model learns from that record, absence becomes a signal that may be mistaken for lower need. Breaking results down by region, age, sex or other relevant groups does not guarantee equity, but it can reveal when an average hides a concentrated failure.
A card that makes responsibility explicit
Before using a prediction, an institution can complete a short card: public-health purpose, population, source and age of data, chosen comparison, types of error, person responsible for reviewing the output, authorised action, complaint mechanism and date of the next evaluation. The WHO guidance places accountability and responsiveness across the full lifecycle rather than in an ethics review detached from deployment.
That card also separates a retrospective test from real validation. Good results on historical data show that a method found regularities in that archive. Testing elsewhere and then observing the system in operation asks whether those regularities survive new populations, procedures and decisions. The greater the possible harm of a wrong alert, the more independent validation should be and the clearer the ability to stop the system.
The transferable skill is specific: when reading a public-health AI claim, identify the decision that follows the prediction and demand the population, denominator, error types, accountable person and review schedule. That reading does not reject technology. It prevents a promising demonstration from being mistaken for a policy ready to care for real people.
One final question connects all the others: which comparison would change the decision? A new system should be tested against existing practice, which may be a simple rule, professional review or no intervention. Comparison with another model alone does not show whether the service improves. Beyond statistical performance, teams should observe time saved, extra workload, health outcomes and unintended effects. If the organisation does not state in advance which result would justify adopting, modifying or withdrawing the tool, evaluation risks becoming a demonstration that can always find a way to look positive. Publishing that decision rule makes uncertainty manageable and gives the public a basis for asking whether the promised benefit appeared.
Sources for this piece
This piece draws on 3 primary source(s), gathered during reporting.
This article was produced with artificial intelligence under human editorial oversight.