Skip Navigation

The Accountability Gap Quietly Killing AI ROI in Financial Services

12 min read • August 2026
Financial institutions are not failing at AI because their models are weak. They are failing because nobody owns the decisions those models now make.

The budget was approved. The pilot worked. The board saw the demonstration and asked when it would reach customers. Then the value disappeared somewhere between proof of concept and production, and no one could say precisely where it went.

The pattern is well documented. Deloitte’s 2025 survey of 1,854 executives found that 85 percent of organizations increased AI investment over the prior year and 91 percent plan to increase it again, yet only 6 percent achieved payback within a single year. Most now expect returns in two to four years rather than the seven to twelve months they originally modelled. RAND Corporation research puts the outright failure rate for AI projects above 80 percent, roughly twice the failure rate of non-AI technology programs.

Read those numbers as a technology problem and the response is predictable: another platform, another model, another pilot. Read them as a governance problem and the response changes entirely. AI systems inherit every characteristic of the environment they operate inside. Inconsistent definitions, unclear lineage and weak access controls are not cleaned up before the model runs. They are encoded directly into what the model produces.

This is the accountability gap. It is not a data team performance issue and it will not be closed by better tooling. The institutions that close it are the ones that decided who owns a decision before they automated it.

Executive Summary

Financial institutions are raising AI budgets faster than they are building the foundations that make AI pay. Deloitte reports 85 percent increasing investment and only 6 percent seeing payback inside a year. The binding constraint is not model quality. It is accountability.

Data teams now sit inside credit, fraud and underwriting decisions, yet they are still measured on pipeline uptime and rarely hold authority to stop a deployment. Agentic AI converts that mismatch into automated, scalable operational risk.

Supervisors have already measured the problem: most firms report only partial understanding of the AI they run. Closing the gap requires semantic governance, decision-grade data products and an operating model with named owners, production gates and decision lineage.

By the numbers

34%

of surveyed UK financial firms report a complete understanding of the AI technologies they use, while 46 percent report only partial understanding

– Bank of England and FCA, 2024

33%

of AI use cases are now third-party implementations, up from 17 percent in 2022, with the top three providers accounting for 73 percent of cloud services

– Bank of England and FCA, 2024

6%

of organizations achieved AI payback within one year, even as 85 percent increased investment and 91 percent planned to increase it again

– Deloitte, 2025

80%+

of AI projects fail outright, roughly twice the failure rate of non-AI technology programs

– RAND Corporation

48%

of AI projects advance past the pilot stage, taking an average of eight months to move from prototype to production

– Gartner, 2024

39%

of organizations attribute any EBIT impact to AI, and most of those say it is under 5 percent, even though 88 percent use AI in at least one function

– McKinsey, 2025

Why Does AI Stall in Banking When the Model Clearly Works?

Scaling AI in a financial institution is not primarily a question of which model to license or which platform to procure. It is a question of whether the data underneath can be trusted to carry a decision that a regulator, a customer or a court may later examine.

AI systems reason across whatever the institution has already built. Where definitions differ between systems, the model does not flag the conflict and escalate it. It resolves the conflict silently and proceeds. A human analyst pauses at a discrepancy. A model treats it as input.

The consequences surface in the same five places in almost every bank.

  • Credit decisions slow down or produce adverse actions nobody can explain, because customer risk is defined differently in origination, servicing and collections.
  • AML false positive rates stay stubbornly high, because surveillance and risk teams do not define suspicious the same way.
  • Customer servicing AI produces inconsistent outcomes, because active customer, eligible product and at-risk account mean different things in the CRM, the core system and the marketing platform.
  • Underwriting models are challenged in examination, because the institution cannot reconstruct which data version, definition and transformation produced a given recommendation.
  • Personalization programs stall, because the business owners who must act on customer data products do not trust them.

None of these are model defects. Every one of them is a governance defect that the model faithfully reproduced at speed.

Supervisors have already quantified how thin the visibility is. In the Bank of England and Financial Conduct Authority survey of artificial intelligence in UK financial services, only 34 percent of firms reported a complete understanding of the AI technologies they use, while 46 percent reported only partial understanding. The regulators attributed much of that gap to reliance on third parties, noting that third-party implementations had risen to a third of all AI use cases from 17 percent in 2022.

Concentration compounds it. The same survey found the top three providers accounted for 73 percent of respondents’ cloud services. When internal understanding is partial and external supply is concentrated, an institution’s ability to explain its own decisions depends on organizations it does not control. That is an accountability problem long before it becomes a technology problem.

The Accountability Shift Nobody Planned For

Data teams were built for a different job. They moved data and prepared it. They were measured on pipeline uptime, dashboard refresh rates and platform availability. The work was infrastructure, and the accountability model matched the work.

AI changed the job without changing the mandate. The pipeline no longer delivers information to a human who then decides. The pipeline participates in the decision. When a model recommends a credit limit, raises a fraud alert or shapes an underwriting outcome, the definitions, versions and lineage behind that model are part of the decision record itself.

So data teams now carry accountability for decision outcomes they were never given the authority to govern. Two questions expose the mismatch immediately, and both should be asked at the next board technology review.

  1. Does the Chief Data Officer have formal standing to halt a model deployment when governance conditions are not met?
  2. Is the data organization measured on decision integrity and explainability, or on pipeline uptime and platform availability?

In most institutions the honest answer to both is no. That is not a data team performance failure. It is a board-level design failure, and design failures are only corrected where the design was made.

Accountability without authority is not governance. It is exposure with a name attached to it.
Key Principle

The organization that is accountable for AI decision integrity must be able to stop an AI deployment.

If it cannot, the institution has not assigned accountability. It has assigned blame in advance, and it has done so to the group with the least authority to prevent the outcome.

How Agentic AI Turns Data Shortcuts Into Operational Risk

Every institution carries data shortcuts: definitions that were never reconciled, lineage that was never captured, access that was granted broadly because narrowing it was slow. Those shortcuts survived because humans absorbed them. An analyst noticed the discrepancy, made a call and moved on.

Agentic AI removes the absorber. It executes the same shortcuts at machine speed, silently, across every workflow it touches, without the pause that used to catch the error.

Failure mode What changes when an agent runs it
Inconsistent definitions become automated errors A human analyst reconciles conflicting definitions before acting. An agent does not. In AML transaction monitoring or credit limit adjustment, the inconsistency is applied at scale with no reconciliation step and no record that a conflict existed
Missing lineage becomes regulatory exposure Regulators expect the institution to show how an AI-influenced decision was reached, including the data behind it. A workflow spanning five data sources and three transformation rules cannot be reconstructed without column-level lineage, and the gap becomes a liability with financial consequences
Access misconfigurations become systemic A misconfigured permission for one analyst affects one analyst. The same error at agent level affects every workflow that agent runs, for as long as it runs. Least privilege has to be enforced at the agent identity level, which traditional identity and access management was not designed to do

The adoption curve is not waiting for the controls. Enterprise Management Associates research in 2025 found that only 2 percent of organizations with more than 500 employees were uninterested in agentic AI, and that 79 percent of organizations without written agentic policies had already deployed agents. Deployment is running ahead of policy, and policy is running ahead of enforcement.

Agentic AI does not create new data risks. It converts the data risks you already have into automated, scalable operational failures.

Why the Regulatory Deadlines Are No Longer Theoretical

Governance debates inside institutions tend to assume time. The supervisory calendar has removed that assumption. The obligations below are not roadmap items. They are live or imminent, and they attach to systems already in production.

Obligation What it means for AI decisions
EU AI Act, Annex III high-risk systems Obligations apply from 2 August 2026, with penalties of up to EUR 35 million or 7 percent of global annual turnover. Credit scoring and several other financial services use cases fall inside the high-risk category
OSFI model risk management guidance Canadian institutions that cannot demonstrate explainability and clear accountability for model-driven decisions should expect examination findings, not extensions
U.S. FSOC 2024 Annual Report Signals increased federal examination scrutiny of AI-driven decision systems and the controls around them
GDPR accountability principle Where an AI agent makes or materially shapes a decision affecting a data subject, the institution must be able to demonstrate how and why that decision was reached

The common requirement across all four is the same, and it is unglamorous. An institution must be able to reconstruct a material AI-influenced decision on demand: the data, the version, the definition, the transformation, the model and the person accountable for it.

The regulatory window has closed on ‘we are working on governance.’ In 2026, institutions that cannot demonstrate explainability, lineage, and defined accountability for AI-influenced decisions are not behind; they are non-compliant.

The inability to reconstruct a material AI-influenced decision is not a gap to be scheduled for remediation in a future cycle. It is a present-tense regulatory liability that grows with every additional workflow the institution automates on the same foundation.

What an AI-Ready Operating Model Actually Requires

Closing the accountability gap is not a platform purchase. It is three connected capabilities, and institutions that skip any one of them rebuild the same problem on newer infrastructure.

1. Semantic governance: defining what every concept means

AI systems do not reason about databases. They reason about concepts. When net revenue, customer lifetime value, net exposure, defaulted account or active relationship are defined differently across origination, risk, servicing and analytics, every model that reasons about a customer operates with structural ambiguity built into its inputs.

The outputs will be inconsistent. And when executives encounter inconsistent AI outputs more than once, they stop trusting the system. Adoption stalls, not because the technology failed, but because the data foundation undermined its credibility before the technology had a chance.

This is the discipline a Semantic Intelligence Framework is built to enforce: definition governance, a lineage protocol and a consumption architecture, tied together so that a single business concept has one approved meaning. Institutions typically start with the 20 to 40 highest-risk concepts rather than attempting the full estate. The target is simple to state and hard to reach: for any concept, the institution can answer instantly what it means, who owns it, when the definition was approved and which models depend on it.

2. Decision-grade data products: what AI systems actually consume

A decision-grade data product is not a dataset. A dataset is a collection of records. A decision-grade data product is an asset designed to be consumed by AI systems and human analysts from the same trusted source, with explicit ownership, documented semantic definitions, defined quality standards, versioning, lineage and contractual expectations for downstream consumers.

The difference is operational, not academic. A credit scoring model consuming a data product knows exactly which version of income it received, who validated it and when it last changed. An AML agent operates inside formal access boundaries rather than inherited ones. A compliance team reconstructs a full provenance chain in hours rather than weeks, or rather than not at all.

3. A governance operating model: who owns what, and what happens when it changes

Strong governance is what allows an agent to run without a human checking every step. Weak governance forces human review into every workflow and still leaves the institution exposed, because the reviews are not evidence.

Required element What it means in practice
Named ownership Every concept and data asset feeding an AI system has one accountable owner. Shared ownership becomes no ownership at the moment a decision is challenged
A production gate No model enters production without approved, versioned definitions for the concepts it reasons about. The gate is a control, not a checklist
Version and lineage protocol A definition change automatically identifies every dependent model and triggers revalidation before those models continue to run
Agent-level access controls Least privilege is enforced at the agent identity level rather than inherited from the human team that commissioned the agent
Decision lineage logging AI-influenced decisions produce a structured audit trail linking the outcome to the data and definitions that shaped it

None of this is free, and leaders should budget for it honestly. McKinsey’s benchmark is that programs allocating 50 to 70 percent of their timeline to data readiness outperform those that do not. Programs that treat data readiness as a preliminary phase to be compressed are the ones that later report a two to four year payback.

Executive Insight

Governance is not what slows AI programs down. The absence of governance is.

Every hour a bank spends reconciling conflicting model outputs, rebuilding lineage for an examination or holding a deployment while ownership is negotiated is an hour the program lost to work that governance would have made unnecessary. The institutions moving fastest are not the least governed. They are the ones whose definitions, owners and gates were settled before the agents started running.

ML arteka works with financial institutions to make that foundation explicit: governed semantics, decision-grade data products and an accountability model that a regulator can follow without a guided tour.

What Leadership Must Do Now

The accountability gap is closed by a small number of decisions taken at the top, not by a large number of initiatives taken below it. The actions below are assignable this quarter.

Role Priority actions
CEO and Board Formally assign accountability for decision integrity to the Chief Data Officer, including authority to halt a deployment. Require a governance readiness gate as a precondition for AI production. Redefine CDO performance measures to include explainability and lineage completeness rather than platform uptime. Commission a semantic risk assessment of the business concepts AI systems reason about
CRO and CCO Identify every AI system in production that lacks full decision lineage traceability. Map current AI deployment against OSFI and EU AI Act obligations and name the gaps. Require an evidence package for each new deployment covering governed definitions, lineage graph, access control documentation and the accountable owner. Move AI decision auditing from periodic to continuous
CFO Require AI ROI reporting to show data readiness cost and timeline as a separate line item. Test budgets against the McKinsey benchmark of 50 to 70 percent of timeline on data readiness. Quantify the cost of ungoverned AI: remediation cycles, compliance rework, AML false positive operations and delayed deployments
COO Assess the five highest-volume AI workflows to establish whether they consume decision-grade data products or merely datasets. Require every agent to document its access boundaries at the agent identity level. Establish a cross-functional escalation path for disputed AI outputs
CDO and CDAO Run a semantic landscape assessment covering the 20 to 40 highest-risk business concepts, mapping definitions, ownership and decision authority. Implement the production gate. Build the lineage graph for the highest-risk flows first: credit, AML and underwriting. Shift team KPIs from pipeline uptime to semantic coverage, lineage completeness and AI decision explainability

Read down that list and one thing is obvious. Almost none of it is technical work. It is the assignment of authority, the definition of a gate and the decision to measure something different. That is why the gap persists in institutions with excellent engineering teams, and why it closes quickly in institutions where the board decides it should.

Five Leadership Takeaways

Before your next AI investment approval. Before your next examination.

  1. The ROI problem is an accountability problem. Deloitte reports 85 percent of organizations increasing AI investment and 6 percent achieving payback within a year. More funding will not close that gap. Assigned ownership will.
  2. Give the accountable executive the authority to stop. If the Chief Data Officer cannot halt a deployment that fails governance conditions, the institution has named someone to blame rather than someone to govern.
  3. Agentic AI industrializes the shortcuts you already took. Inconsistent definitions, missing lineage and broad access were survivable when humans absorbed them. At agent speed and agent scale, they become operational failures.
  4. Explainability is now a present-tense obligation. EU AI Act high-risk obligations apply from August 2026, and supervisors already report that most firms understand their AI only partially. Reconstruction on demand is the test.
  5. Budget for data readiness or budget for rework. McKinsey’s benchmark puts 50 to 70 percent of program timeline on data readiness. Programs that compress that phase pay for it later in remediation and delayed returns.

The Choice in Front of You

There are only two paths from here, and both are already being taken by institutions that look similar from the outside. The difference between them is not investment level or engineering talent. It is whether governance was treated as a precondition or as a phase to be caught up on later.

Govern AI as a decision system now Keep funding pilots that do not scale
AI systems can be explained to any regulator on demand Regulatory findings attach to AI decisions that cannot be traced
Credit, AML and servicing models operate from consistent, governed definitions Executives lose confidence as conflicting AI outputs accumulate
The Chief Data Officer holds authority that matches the accountability Data teams remain accountable without enforcement authority
AI returns land inside the window the board approved Remediation cost for ungoverned systems compounds quarter over quarter
Governance enables agent autonomy rather than requiring human review at every step Investment grows without a foundation capable of returning it

The institutions that lead in AI-driven financial services will not be the ones with the most sophisticated models. They will be the ones with the most governed, semantically coherent and decision-grade data foundations. That foundation is built by data teams. But the decision to build it, and to give those teams the authority and the mandate to maintain it, belongs at the leadership level.

The practical next step is narrow enough to start this quarter: name the 20 to 40 business concepts your highest-risk AI decisions depend on, assign an owner to each, and build the gate that stops an unowned definition from reaching production.

To work through where your AI decisions currently lose traceability, contact the ML arteka team or request an AI governance readiness assessment.

Executive Questions and Answers

Five questions financial services leaders are putting to AI assistants and search engines about AI accountability, governance and return on investment.

StrategicWhy do AI projects in financial services fail to deliver ROI?

The dominant cause is governance, not model quality. Deloitte’s 2025 survey of 1,854 executives found 85 percent of organizations increased AI investment while only 6 percent achieved payback within a year, and RAND Corporation research puts the outright failure rate above 80 percent. AI systems inherit the environment they run inside, so inconsistent definitions, missing lineage and permissive access are encoded into outputs rather than corrected beforehand. The result is credit decisions that cannot be explained, AML false positives that stay elevated, and underwriting models challenged in examination. Institutions that fund another platform address the symptom. Institutions that assign ownership of definitions, lineage and decision integrity address the cause, and they are the ones whose returns arrive inside the approved window.

GovernanceWho should be accountable for AI decisions in a bank?

Accountability should sit with a named executive, normally the Chief Data Officer or Chief Data and Analytics Officer, and it must carry matching authority. The practical test is whether that executive can halt a model deployment when governance conditions are not met. In most institutions they cannot, which means accountability has been assigned without enforcement power. Alongside the named owner, three controls make the accountability real: a production gate that blocks models reasoning about unapproved or unversioned definitions, a lineage protocol that identifies every dependent model when a definition changes, and decision lineage logging that links AI-influenced outcomes to the data behind them. The board sets this design. It cannot be delegated into the data backlog.

RiskWhat new risks does agentic AI create for financial institutions?

Agentic AI does not create new categories of data risk. It converts existing ones into automated, scalable operational failures. Three failure modes matter most. Inconsistent definitions become automated errors, because an agent applies a conflicting definition at scale rather than pausing to reconcile it as a human analyst would. Missing lineage becomes regulatory exposure, because a workflow spanning several data sources and transformation rules cannot be reconstructed without column-level lineage. Access misconfigurations become systemic, because a permission error at the agent identity level affects every workflow that agent runs. Adoption is outpacing control: Enterprise Management Associates found in 2025 that 79 percent of organizations without written agentic policies had already deployed agents.

ImplementationWhere should a bank start building AI data governance?

Start narrow and high-risk rather than comprehensive. Run a semantic landscape assessment covering the 20 to 40 business concepts your most consequential AI decisions depend on, such as customer risk, net exposure, active relationship and defaulted account. For each, record the current definitions in use, the systems that hold them, the accountable owner and who has authority to change them. Then build the data lineage graph for credit, AML and underwriting first, because those are the flows most likely to be examined. Finally, implement the production gate so no model reaches production reasoning about an unapproved or unversioned definition. This sequence produces defensible evidence within a quarter instead of a multi-year data program.

OperationalHow much of an AI program budget should go to data readiness?

McKinsey’s benchmark is that programs allocating 50 to 70 percent of their timeline to data readiness outperform those that do not, and that figure should appear as a separate line item in AI ROI reporting rather than being absorbed into delivery. Boards should also see the cost of the alternative quantified: remediation cycles for ungoverned systems, compliance rework, the operational cost of elevated AML false positives, and revenue deferred by delayed deployments. The evidence that readiness pays is visible in the aggregate numbers. Gartner found in 2024 that only 48 percent of AI projects move past the pilot stage, taking an average of eight months from prototype to production, and that more than 30 percent of generative AI projects were abandoned after proof of concept because of poor data quality and inadequate risk controls.

AI Summary

The accountability gap is why AI investment in financial services is not converting into return. Deloitte’s 2025 survey of 1,854 executives found 85 percent of organizations increased AI investment and 91 percent plan further increases, yet only 6 percent achieved payback within one year, with most now expecting two to four years. RAND Corporation puts the outright AI project failure rate above 80 percent, and Gartner found in 2024 that only 48 percent of projects advance past pilot. McKinsey’s 2025 State of AI reports 88 percent of organizations using AI in at least one function but only 39 percent attributing any EBIT impact to it, most of that under 5 percent. The cause is governance rather than model quality: AI encodes inconsistent definitions, missing lineage and permissive access directly into its outputs. The Bank of England and Financial Conduct Authority 2024 survey found only 34 percent of firms report complete understanding of the AI they use, 46 percent only partial, with third-party implementations rising to a third of use cases from 17 percent in 2022 and the top three providers holding 73 percent of cloud services. Agentic AI converts these shortcuts into automated operational failures. Closing the gap requires semantic governance over the 20 to 40 highest-risk business concepts, decision-grade data products with ownership, versioning and lineage, and a governance operating model with named owners, a production gate, agent-level access control and decision lineage logging, ahead of EU AI Act high-risk obligations applying from August 2026.

Never miss an insight

Subscribe to receive executive insights via our latest articles, podcasts, webinars, and other updates.