AI in Finance: What Exists and What Remains Hypothetical AGI
AI already supports fraud detection, analysis and financial operations, but that is not evidence of AGI. A guide to separating capability, autonomy and accountability.
On 21 November 2024, the Bank of England and the Financial Conduct Authority published a survey of 118 regulated firms: 75% said they were already using artificial intelligence and another 10% planned to do so within three years. Foundation models accounted for 17% of reported use cases. Those are figures about AI adoption, not evidence of artificial general intelligence (AGI). The distinction may sound terminological, but it changes what a financial firm can promise, what it must measure and who is accountable when a system fails.
The useful question is not whether a bank, insurer or fund “uses AGI”. It is more concrete: what task does the system perform, how reliable is it, how much autonomy does it have, and what controls surround its decisions? Asked this way, the language no longer presents a future capability as though it were already deployed.
Current AI and AGI are not two versions of the same product
A system can outperform every person at one narrow task and still be specialised AI. The Levels of AGI framework, developed by Google DeepMind researchers, separates two dimensions: performance —the depth of capability— and generality —the breadth of tasks on which that performance is achieved. Its table places AlphaZero under “superhuman narrow AI”: it masters specific games, not the full range of cognitive work performed by a financial professional. The paper also says that “competent AGI”, able to perform at the level of a skilled adult across most cognitive tasks, had not been achieved by any public system at the time assessed.
This distinction corrects two common shortcuts. First, a language model that writes, summarises or extracts information does not become AGI because it participates in many workflows. It may be a general-purpose tool with uneven performance. Second, success in chess, Go or simulation does not show that a system can transfer that skill on its own to lending, compliance or portfolio management. Transfer has to be measured; it cannot be inherited by analogy.
There is also insufficient public evidence for the claim that Bridgewater or Renaissance Technologies are “investing in AGI”. They may employ statistics, machine learning or other quantitative methods, but that does not justify moving those methods into a different category. When a company does not publish the system, the data or the test, the honest description is “not verifiable”, not “secret AGI”.
What is actually happening in finance
Deployed AI already creates value without being called general. The Bank of England and FCA survey found the greatest perceived current benefits in data and analytical insights, anti-money laundering and fraud prevention, and cybersecurity. Operations and IT accounted for about 22% of reported use cases. The Financial Stability Board also identifies uses in credit assessment, customer interaction, capital optimisation, model-risk management, trading, portfolio management, regulatory technology and supervision.
These are different uses and must be assessed separately. Classifying a suspicious transaction is not the same as deciding whether a person receives a loan. Summarising a file for an analyst is not the same as signing off the transaction. Producing an alert is not the same as executing a market order. The word “AI” describes a family of techniques; on its own, it does not reveal the authority given to the system.
The UK survey itself provides an antidote to adoption headlines. Firms classified 62% of all use cases as low materiality and 16% as high materiality. In addition, 46% of firms reported only a “partial understanding” of the AI technology they used, compared with 34% that reported a “complete understanding”. Knowing that an institution “uses AI” tells us little unless we know the task, the impact and the degree of dependency.
Four questions for reading any “financial AGI” claim
1. What breadth has been demonstrated? Ask for the list of evaluated tasks, not a striking demo. A system that analyses documents, generates code and holds a conversation may look broad; it must still be tested on whether it learns new tasks, recognises when it needs help and maintains performance outside prepared examples.
2. What does “better” mean? A laboratory result may measure average accuracy, simulated returns or response time. None of those alone demonstrates robustness to a market shift, incomplete data or a crisis. The comparison must stay on the same axis: an improvement in text extraction is not evidence of better credit decisions.
3. How autonomous is it? Suggesting, ranking, executing and learning without approval are not equivalent. Risk depends on permission as well as capability. A powerful model inside a workflow with limits, human review and reversibility may pose less risk than a more modest model authorised to act at high speed.
4. Who is accountable, and how can the decision be reversed? There should be an identifiable owner, a record of inputs and outputs, an appeal process, a fallback system and a criterion for stopping the model. If those answers are missing, debating whether the tool is “general” distracts from the immediate problem.
The risk map already exists
The principal risks do not require AGI. They begin with data and models: incomplete information, selection bias, distribution shifts, hard-to-explain outputs and convincing errors from generative models. In the UK survey, four of the five most-cited current risks involved data: privacy and protection, quality, security, and bias or representativeness.
The second layer concerns decisions about people. The EU AI Act classifies certain systems intended to assess a person’s creditworthiness or credit score, and to assess risk or set prices in life and health insurance, as high-risk. It also distinguishes those uses from systems used to detect fraud in financial services or to perform prudential calculations provided for in EU law. The durable lesson is that risk is assigned by purpose and effect, not by the model’s brand name.
The third layer is shared dependency. One third of the use cases in the UK survey were third-party implementations; the three largest providers accounted for 73% of reported cloud services and 44% of models. The Financial Stability Board identifies service-provider concentration, market correlations, cyber risk, and model, data and governance risk as vulnerabilities that can amplify systemic risk. Institutions may appear diversified while relying on the same infrastructure or similar models.
The fourth layer is adversarial. On 27 March 2024, the US Treasury Department documented capability gaps between large and small institutions, a shortage of shared data for fighting fraud, explainability problems and the need to map the data supply chain. A financial system must do more than work on ordinary data: it must withstand manipulation, impersonation and inputs designed to deceive it.
Govern the real system, not the label
A serious assessment can be organised around the four functions of the NIST AI Risk Management Framework: govern, map, measure and manage. In finance, that means assigning an owner and limits; describing who is affected by the decision; measuring performance, bias and security in relevant scenarios; and monitoring the system after deployment. The process is continuous because markets, data, attacks and suppliers change.
Before approving a use case, an institution should require, at minimum, data provenance, a comparison with a simpler alternative, out-of-sample and stress testing, autonomy limits, effective human review, an auditable record, drift monitoring, an incident plan and a supplier exit route. None of these controls depends on AGI arriving. If a future system genuinely expands its generality, the controls will need to become stronger, not disappear.
The reader’s transferable skill is concrete: when faced with a “financial AGI” claim, separate breadth, performance, autonomy and accountability. If the evidence covers only one task, one benchmark or one demonstration, it describes a useful but bounded capability. Naming it precisely does not diminish its value; it prevents a hypothesis about the future from masquerading as a description of the present.
This article was produced with artificial intelligence under human editorial oversight.