The budget was approved. The pilot worked. The board saw the demonstration and asked when it would reach customers. Then the value disappeared somewhere between proof of concept and production, and no one could say precisely where it went.
The pattern is well documented. Deloitte’s 2025 survey of 1,854 executives found that 85 percent of organizations increased AI investment over the prior year and 91 percent plan to increase it again, yet only 6 percent achieved payback within a single year. Most now expect returns in two to four years rather than the seven to twelve months they originally modelled. RAND Corporation research puts the outright failure rate for AI projects above 80 percent, roughly twice the failure rate of non-AI technology programs.
Read those numbers as a technology problem and the response is predictable: another platform, another model, another pilot. Read them as a governance problem and the response changes entirely. AI systems inherit every characteristic of the environment they operate inside. Inconsistent definitions, unclear lineage and weak access controls are not cleaned up before the model runs. They are encoded directly into what the model produces.
This is the accountability gap. It is not a data team performance issue and it will not be closed by better tooling. The institutions that close it are the ones that decided who owns a decision before they automated it.
Financial institutions are raising AI budgets faster than they are building the foundations that make AI pay. Deloitte reports 85 percent increasing investment and only 6 percent seeing payback inside a year. The binding constraint is not model quality. It is accountability.
Data teams now sit inside credit, fraud and underwriting decisions, yet they are still measured on pipeline uptime and rarely hold authority to stop a deployment. Agentic AI converts that mismatch into automated, scalable operational risk.
Supervisors have already measured the problem: most firms report only partial understanding of the AI they run. Closing the gap requires semantic governance, decision-grade data products and an operating model with named owners, production gates and decision lineage.
By the numbers
34%
of surveyed UK financial firms report a complete understanding of the AI technologies they use, while 46 percent report only partial understanding
– Bank of England and FCA, 2024
33%
of AI use cases are now third-party implementations, up from 17 percent in 2022, with the top three providers accounting for 73 percent of cloud services
– Bank of England and FCA, 2024
6%
of organizations achieved AI payback within one year, even as 85 percent increased investment and 91 percent planned to increase it again
– Deloitte, 2025
80%+
of AI projects fail outright, roughly twice the failure rate of non-AI technology programs
– RAND Corporation
48%
of AI projects advance past the pilot stage, taking an average of eight months to move from prototype to production
– Gartner, 2024
39%
of organizations attribute any EBIT impact to AI, and most of those say it is under 5 percent, even though 88 percent use AI in at least one function
– McKinsey, 2025
Why Does AI Stall in Banking When the Model Clearly Works?
Scaling AI in a financial institution is not primarily a question of which model to license or which platform to procure. It is a question of whether the data underneath can be trusted to carry a decision that a regulator, a customer or a court may later examine.
AI systems reason across whatever the institution has already built. Where definitions differ between systems, the model does not flag the conflict and escalate it. It resolves the conflict silently and proceeds. A human analyst pauses at a discrepancy. A model treats it as input.
The consequences surface in the same five places in almost every bank.
- Credit decisions slow down or produce adverse actions nobody can explain, because customer risk is defined differently in origination, servicing and collections.
- AML false positive rates stay stubbornly high, because surveillance and risk teams do not define suspicious the same way.
- Customer servicing AI produces inconsistent outcomes, because active customer, eligible product and at-risk account mean different things in the CRM, the core system and the marketing platform.
- Underwriting models are challenged in examination, because the institution cannot reconstruct which data version, definition and transformation produced a given recommendation.
- Personalization programs stall, because the business owners who must act on customer data products do not trust them.
None of these are model defects. Every one of them is a governance defect that the model faithfully reproduced at speed.
Supervisors have already quantified how thin the visibility is. In the Bank of England and Financial Conduct Authority survey of artificial intelligence in UK financial services, only 34 percent of firms reported a complete understanding of the AI technologies they use, while 46 percent reported only partial understanding. The regulators attributed much of that gap to reliance on third parties, noting that third-party implementations had risen to a third of all AI use cases from 17 percent in 2022.
Concentration compounds it. The same survey found the top three providers accounted for 73 percent of respondents’ cloud services. When internal understanding is partial and external supply is concentrated, an institution’s ability to explain its own decisions depends on organizations it does not control. That is an accountability problem long before it becomes a technology problem.
Related reading: 2026 Trends Guiding Financial Services Out of AI Pilot Purgatory
The Accountability Shift Nobody Planned For
Data teams were built for a different job. They moved data and prepared it. They were measured on pipeline uptime, dashboard refresh rates and platform availability. The work was infrastructure, and the accountability model matched the work.
AI changed the job without changing the mandate. The pipeline no longer delivers information to a human who then decides. The pipeline participates in the decision. When a model recommends a credit limit, raises a fraud alert or shapes an underwriting outcome, the definitions, versions and lineage behind that model are part of the decision record itself.
So data teams now carry accountability for decision outcomes they were never given the authority to govern. Two questions expose the mismatch immediately, and both should be asked at the next board technology review.
- Does the Chief Data Officer have formal standing to halt a model deployment when governance conditions are not met?
- Is the data organization measured on decision integrity and explainability, or on pipeline uptime and platform availability?
In most institutions the honest answer to both is no. That is not a data team performance failure. It is a board-level design failure, and design failures are only corrected where the design was made.
The organization that is accountable for AI decision integrity must be able to stop an AI deployment.
If it cannot, the institution has not assigned accountability. It has assigned blame in advance, and it has done so to the group with the least authority to prevent the outcome.
How Agentic AI Turns Data Shortcuts Into Operational Risk
Every institution carries data shortcuts: definitions that were never reconciled, lineage that was never captured, access that was granted broadly because narrowing it was slow. Those shortcuts survived because humans absorbed them. An analyst noticed the discrepancy, made a call and moved on.
Agentic AI removes the absorber. It executes the same shortcuts at machine speed, silently, across every workflow it touches, without the pause that used to catch the error.
| Failure mode | What changes when an agent runs it |
|---|---|
| Inconsistent definitions become automated errors | A human analyst reconciles conflicting definitions before acting. An agent does not. In AML transaction monitoring or credit limit adjustment, the inconsistency is applied at scale with no reconciliation step and no record that a conflict existed |
| Missing lineage becomes regulatory exposure | Regulators expect the institution to show how an AI-influenced decision was reached, including the data behind it. A workflow spanning five data sources and three transformation rules cannot be reconstructed without column-level lineage, and the gap becomes a liability with financial consequences |
| Access misconfigurations become systemic | A misconfigured permission for one analyst affects one analyst. The same error at agent level affects every workflow that agent runs, for as long as it runs. Least privilege has to be enforced at the agent identity level, which traditional identity and access management was not designed to do |
The adoption curve is not waiting for the controls. Enterprise Management Associates research in 2025 found that only 2 percent of organizations with more than 500 employees were uninterested in agentic AI, and that 79 percent of organizations without written agentic policies had already deployed agents. Deployment is running ahead of policy, and policy is running ahead of enforcement.
Why the Regulatory Deadlines Are No Longer Theoretical
Governance debates inside institutions tend to assume time. The supervisory calendar has removed that assumption. The obligations below are not roadmap items. They are live or imminent, and they attach to systems already in production.
| Obligation | What it means for AI decisions |
|---|---|
| EU AI Act, Annex III high-risk systems | Obligations apply from 2 August 2026, with penalties of up to EUR 35 million or 7 percent of global annual turnover. Credit scoring and several other financial services use cases fall inside the high-risk category |
| OSFI model risk management guidance | Canadian institutions that cannot demonstrate explainability and clear accountability for model-driven decisions should expect examination findings, not extensions |
| U.S. FSOC 2024 Annual Report | Signals increased federal examination scrutiny of AI-driven decision systems and the controls around them |
| GDPR accountability principle | Where an AI agent makes or materially shapes a decision affecting a data subject, the institution must be able to demonstrate how and why that decision was reached |
The common requirement across all four is the same, and it is unglamorous. An institution must be able to reconstruct a material AI-influenced decision on demand: the data, the version, the definition, the transformation, the model and the person accountable for it.
The inability to reconstruct a material AI-influenced decision is not a gap to be scheduled for remediation in a future cycle. It is a present-tense regulatory liability that grows with every additional workflow the institution automates on the same foundation.
What an AI-Ready Operating Model Actually Requires
Closing the accountability gap is not a platform purchase. It is three connected capabilities, and institutions that skip any one of them rebuild the same problem on newer infrastructure.
1. Semantic governance: defining what every concept means
AI systems do not reason about databases. They reason about concepts. When net revenue, customer lifetime value, net exposure, defaulted account or active relationship are defined differently across origination, risk, servicing and analytics, every model that reasons about a customer operates with structural ambiguity built into its inputs.
The outputs will be inconsistent. And when executives encounter inconsistent AI outputs more than once, they stop trusting the system. Adoption stalls, not because the technology failed, but because the data foundation undermined its credibility before the technology had a chance.
This is the discipline a Semantic Intelligence Framework is built to enforce: definition governance, a lineage protocol and a consumption architecture, tied together so that a single business concept has one approved meaning. Institutions typically start with the 20 to 40 highest-risk concepts rather than attempting the full estate. The target is simple to state and hard to reach: for any concept, the institution can answer instantly what it means, who owns it, when the definition was approved and which models depend on it.
2. Decision-grade data products: what AI systems actually consume
A decision-grade data product is not a dataset. A dataset is a collection of records. A decision-grade data product is an asset designed to be consumed by AI systems and human analysts from the same trusted source, with explicit ownership, documented semantic definitions, defined quality standards, versioning, lineage and contractual expectations for downstream consumers.
The difference is operational, not academic. A credit scoring model consuming a data product knows exactly which version of income it received, who validated it and when it last changed. An AML agent operates inside formal access boundaries rather than inherited ones. A compliance team reconstructs a full provenance chain in hours rather than weeks, or rather than not at all.
3. A governance operating model: who owns what, and what happens when it changes
Strong governance is what allows an agent to run without a human checking every step. Weak governance forces human review into every workflow and still leaves the institution exposed, because the reviews are not evidence.
| Required element | What it means in practice |
|---|---|
| Named ownership | Every concept and data asset feeding an AI system has one accountable owner. Shared ownership becomes no ownership at the moment a decision is challenged |
| A production gate | No model enters production without approved, versioned definitions for the concepts it reasons about. The gate is a control, not a checklist |
| Version and lineage protocol | A definition change automatically identifies every dependent model and triggers revalidation before those models continue to run |
| Agent-level access controls | Least privilege is enforced at the agent identity level rather than inherited from the human team that commissioned the agent |
| Decision lineage logging | AI-influenced decisions produce a structured audit trail linking the outcome to the data and definitions that shaped it |
None of this is free, and leaders should budget for it honestly. McKinsey’s benchmark is that programs allocating 50 to 70 percent of their timeline to data readiness outperform those that do not. Programs that treat data readiness as a preliminary phase to be compressed are the ones that later report a two to four year payback.
Governance is not what slows AI programs down. The absence of governance is.
Every hour a bank spends reconciling conflicting model outputs, rebuilding lineage for an examination or holding a deployment while ownership is negotiated is an hour the program lost to work that governance would have made unnecessary. The institutions moving fastest are not the least governed. They are the ones whose definitions, owners and gates were settled before the agents started running.
ML arteka works with financial institutions to make that foundation explicit: governed semantics, decision-grade data products and an accountability model that a regulator can follow without a guided tour.
What Leadership Must Do Now
The accountability gap is closed by a small number of decisions taken at the top, not by a large number of initiatives taken below it. The actions below are assignable this quarter.
| Role | Priority actions |
|---|---|
| CEO and Board | Formally assign accountability for decision integrity to the Chief Data Officer, including authority to halt a deployment. Require a governance readiness gate as a precondition for AI production. Redefine CDO performance measures to include explainability and lineage completeness rather than platform uptime. Commission a semantic risk assessment of the business concepts AI systems reason about |
| CRO and CCO | Identify every AI system in production that lacks full decision lineage traceability. Map current AI deployment against OSFI and EU AI Act obligations and name the gaps. Require an evidence package for each new deployment covering governed definitions, lineage graph, access control documentation and the accountable owner. Move AI decision auditing from periodic to continuous |
| CFO | Require AI ROI reporting to show data readiness cost and timeline as a separate line item. Test budgets against the McKinsey benchmark of 50 to 70 percent of timeline on data readiness. Quantify the cost of ungoverned AI: remediation cycles, compliance rework, AML false positive operations and delayed deployments |
| COO | Assess the five highest-volume AI workflows to establish whether they consume decision-grade data products or merely datasets. Require every agent to document its access boundaries at the agent identity level. Establish a cross-functional escalation path for disputed AI outputs |
| CDO and CDAO | Run a semantic landscape assessment covering the 20 to 40 highest-risk business concepts, mapping definitions, ownership and decision authority. Implement the production gate. Build the lineage graph for the highest-risk flows first: credit, AML and underwriting. Shift team KPIs from pipeline uptime to semantic coverage, lineage completeness and AI decision explainability |
Read down that list and one thing is obvious. Almost none of it is technical work. It is the assignment of authority, the definition of a gate and the decision to measure something different. That is why the gap persists in institutions with excellent engineering teams, and why it closes quickly in institutions where the board decides it should.
Five Leadership Takeaways
Before your next AI investment approval. Before your next examination.
- 1The ROI problem is an accountability problem. Deloitte reports 85 percent of organizations increasing AI investment and 6 percent achieving payback within a year. More funding will not close that gap. Assigned ownership will.
- 2Give the accountable executive the authority to stop. If the Chief Data Officer cannot halt a deployment that fails governance conditions, the institution has named someone to blame rather than someone to govern.
- 3Agentic AI industrializes the shortcuts you already took. Inconsistent definitions, missing lineage and broad access were survivable when humans absorbed them. At agent speed and agent scale, they become operational failures.
- 4Explainability is now a present-tense obligation. EU AI Act high-risk obligations apply from August 2026, and supervisors already report that most firms understand their AI only partially. Reconstruction on demand is the test.
- 5Budget for data readiness or budget for rework. McKinsey’s benchmark puts 50 to 70 percent of program timeline on data readiness. Programs that compress that phase pay for it later in remediation and delayed returns.
The Choice in Front of You
There are only two paths from here, and both are already being taken by institutions that look similar from the outside. The difference between them is not investment level or engineering talent. It is whether governance was treated as a precondition or as a phase to be caught up on later.
| Govern AI as a decision system now | Keep funding pilots that do not scale |
|---|---|
| AI systems can be explained to any regulator on demand | Regulatory findings attach to AI decisions that cannot be traced |
| Credit, AML and servicing models operate from consistent, governed definitions | Executives lose confidence as conflicting AI outputs accumulate |
| The Chief Data Officer holds authority that matches the accountability | Data teams remain accountable without enforcement authority |
| AI returns land inside the window the board approved | Remediation cost for ungoverned systems compounds quarter over quarter |
| Governance enables agent autonomy rather than requiring human review at every step | Investment grows without a foundation capable of returning it |
The institutions that lead in AI-driven financial services will not be the ones with the most sophisticated models. They will be the ones with the most governed, semantically coherent and decision-grade data foundations. That foundation is built by data teams. But the decision to build it, and to give those teams the authority and the mandate to maintain it, belongs at the leadership level.
The practical next step is narrow enough to start this quarter: name the 20 to 40 business concepts your highest-risk AI decisions depend on, assign an owner to each, and build the gate that stops an unowned definition from reaching production.
To work through where your AI decisions currently lose traceability, contact the ML arteka team or request an AI governance readiness assessment.
Executive Questions and Answers
Five questions financial services leaders are putting to AI assistants and search engines about AI accountability, governance and return on investment.
StrategicWhy do AI projects in financial services fail to deliver ROI?
The dominant cause is governance, not model quality. Deloitte’s 2025 survey of 1,854 executives found 85 percent of organizations increased AI investment while only 6 percent achieved payback within a year, and RAND Corporation research puts the outright failure rate above 80 percent. AI systems inherit the environment they run inside, so inconsistent definitions, missing lineage and permissive access are encoded into outputs rather than corrected beforehand. The result is credit decisions that cannot be explained, AML false positives that stay elevated, and underwriting models challenged in examination. Institutions that fund another platform address the symptom. Institutions that assign ownership of definitions, lineage and decision integrity address the cause, and they are the ones whose returns arrive inside the approved window.
GovernanceWho should be accountable for AI decisions in a bank?
Accountability should sit with a named executive, normally the Chief Data Officer or Chief Data and Analytics Officer, and it must carry matching authority. The practical test is whether that executive can halt a model deployment when governance conditions are not met. In most institutions they cannot, which means accountability has been assigned without enforcement power. Alongside the named owner, three controls make the accountability real: a production gate that blocks models reasoning about unapproved or unversioned definitions, a lineage protocol that identifies every dependent model when a definition changes, and decision lineage logging that links AI-influenced outcomes to the data behind them. The board sets this design. It cannot be delegated into the data backlog.
RiskWhat new risks does agentic AI create for financial institutions?
Agentic AI does not create new categories of data risk. It converts existing ones into automated, scalable operational failures. Three failure modes matter most. Inconsistent definitions become automated errors, because an agent applies a conflicting definition at scale rather than pausing to reconcile it as a human analyst would. Missing lineage becomes regulatory exposure, because a workflow spanning several data sources and transformation rules cannot be reconstructed without column-level lineage. Access misconfigurations become systemic, because a permission error at the agent identity level affects every workflow that agent runs. Adoption is outpacing control: Enterprise Management Associates found in 2025 that 79 percent of organizations without written agentic policies had already deployed agents.
ImplementationWhere should a bank start building AI data governance?
Start narrow and high-risk rather than comprehensive. Run a semantic landscape assessment covering the 20 to 40 business concepts your most consequential AI decisions depend on, such as customer risk, net exposure, active relationship and defaulted account. For each, record the current definitions in use, the systems that hold them, the accountable owner and who has authority to change them. Then build the data lineage graph for credit, AML and underwriting first, because those are the flows most likely to be examined. Finally, implement the production gate so no model reaches production reasoning about an unapproved or unversioned definition. This sequence produces defensible evidence within a quarter instead of a multi-year data program.
OperationalHow much of an AI program budget should go to data readiness?
McKinsey’s benchmark is that programs allocating 50 to 70 percent of their timeline to data readiness outperform those that do not, and that figure should appear as a separate line item in AI ROI reporting rather than being absorbed into delivery. Boards should also see the cost of the alternative quantified: remediation cycles for ungoverned systems, compliance rework, the operational cost of elevated AML false positives, and revenue deferred by delayed deployments. The evidence that readiness pays is visible in the aggregate numbers. Gartner found in 2024 that only 48 percent of AI projects move past the pilot stage, taking an average of eight months from prototype to production, and that more than 30 percent of generative AI projects were abandoned after proof of concept because of poor data quality and inadequate risk controls.
Related Content
Articles
The accountability gap is why AI investment in financial services is not converting into return. Deloitte’s 2025 survey of 1,854 executives found 85 percent of organizations increased AI investment and 91 percent plan further increases, yet only 6 percent achieved payback within one year, with most now expecting two to four years. RAND Corporation puts the outright AI project failure rate above 80 percent, and Gartner found in 2024 that only 48 percent of projects advance past pilot. McKinsey’s 2025 State of AI reports 88 percent of organizations using AI in at least one function but only 39 percent attributing any EBIT impact to it, most of that under 5 percent. The cause is governance rather than model quality: AI encodes inconsistent definitions, missing lineage and permissive access directly into its outputs. The Bank of England and Financial Conduct Authority 2024 survey found only 34 percent of firms report complete understanding of the AI they use, 46 percent only partial, with third-party implementations rising to a third of use cases from 17 percent in 2022 and the top three providers holding 73 percent of cloud services. Agentic AI converts these shortcuts into automated operational failures. Closing the gap requires semantic governance over the 20 to 40 highest-risk business concepts, decision-grade data products with ownership, versioning and lineage, and a governance operating model with named owners, a production gate, agent-level access control and decision lineage logging, ahead of EU AI Act high-risk obligations applying from August 2026.