For the past decade, I’ve led customer data and regulatory platforms at one of the world’s largest card issuers, following years modernizing mainframe systems at IBM.

The mainframes I maintained early in my career were not sophisticated by modern standards. They earned the trust of the organizations that depended on them anyway, because every transaction was traceable, every failure was diagnosable, and every change was documented. The agentic AI systems financial institutions are building today need to meet that same standard, and twenty years of new technology has not changed what it takes to earn it.

Alexis Peter Francis

I spent those early years on systems that nobody wanted to touch: COBOL programs were inherited from engineers who had long since retired, and batch jobs that ran overnight and, if they failed, compromised the entire next business day. The knowledge required to troubleshoot these systems existed primarily in the heads of a shrinking pool of specialists, and the stakes were high enough that wholesale migrations off mainframes remain exceedingly rare, with only a 0.2 percent replacement rate in their survey of mainframe shops.

The lesson took hold during large-scale modernization work at IBM, where we exposed COBOL and CICS transaction systems through web services without disrupting overnight batch processing. We could not simply modernize and hope for the best. Every interface had to preserve data lineage, error handling, and rollback behavior because downstream systems depended on deterministic outcomes.

Years later, when agentic AI entered the conversation, I recognized the same pattern. Deploying it without the same discipline around traceability and failure modes.

Financial institutions face a similar challenge with AI adoption. 75 percent of firms already using AI, with measurable productivity gains nearly quadrupling in AI-exposed industries since 2022. But the same survey revealed that only 4.5 percent of organizations trust AI to act fully autonomously. Nearly half require AI systems to make recommendations while reserving final decision-making authority for humans.

When I led implementation of agentic AI for counterparty hierarchy validation at American Express, the goal was never to remove humans from the process. The goal was to reduce manual review effort by 50 percent while preserving full human attestation and audit traceability. In highly regulated domains such as SCCL exposure aggregation and FDIC 370 reporting, defensibility is the true measure of success. Fully autonomous AI decisions, if later found to contain errors in counterparty hierarchies or exposure calculations, can trigger significant regulatory fines, enforcement actions, mandated remediation programs, and in extreme cases, restrictions on business activities or orders to cease certain operations. Several peer banks have faced supervisory scrutiny and reputational damage due to inaccurate regulatory reporting driven by flawed data aggregation. I proactively studied these incidents and incorporated those lessons into our governance framework to prevent similar outcomes.

Through targeted case studies, I demonstrated how seemingly minor inaccuracies in counterparty hierarchies could cascade into materially misstated credit exposure assessments. Systems generating regulatory metrics must withstand audit scrutiny years later, with clear data lineage, explainability, and accountable ownership. By embedding structured human attestation and expert review into the workflow, we preserved accountability, strengthened exception management, and ensured every material output could be defended to regulators. This disciplined hybrid model ultimately accelerated enterprise adoption by building trust across risk, compliance, audit, and business stakeholders.

An earlier experience shaped this instinct. During the Express Scripts and Medco merger  I led large-scale data modernization that integrated millions of member records, eligibility rules, historical claims, and provider relationships under strict operational timelines. One near-miss involved reference data inconsistencies between the two organizationsโ€™ eligibility systems. While data appeared structurally aligned, subtle differences in how legacy platforms encoded relationship hierarchies and benefit qualifiers would have caused downstream adjudication errors if migrated without deeper validation. The risk was not missing data; it was misinterpreting embedded business logic accumulated over decades.

Customer data platforms present an especially instructive case. The work of building a unified view across 248 million customers, 500 million relationships, and 3.5 billion network linkages requires the same obsessive attention to data lineage that characterized mainframe environments. Every batch job in a legacy system had clear inputs, outputs, and failure modes. Every customer record in a modern Customer 360 platform needs the same transparency. The biggest challenges for implementing Customer 360 remain fragmentation and inconsistency of master data, often stored across disparate systems without shared formats or definitions. Entity resolution at this scale requires proprietary algorithms precisely because vendor tools cannot accommodate the institutional knowledge embedded in how data was originally structured.

Iโ€™ve had direct conversations with auditors and regulators around AI-enabled and data-driven decision systems. The questions are remarkably consistent: How was the decision made? What data was used? Who approved it? And can the outcome be reproduced? What reassures them is not the sophistication of the model, but the clarity of the controls. When we can show evidence lineage, human approval checkpoints, and deterministic audit logs, the conversation shifts from skepticism to confidence. Research confirms that human-in-the-loop oversight is becoming foundational for operationalizing trust, not a safety net bolted on after deployment.

Regulators are starting to codify these principles. The FCA launched its AI lab in October 2024 with initiatives including a Supercharged Sandbox for firms to experiment with AI in a secure environment, and research into AI bias that has already informed public discussion on fairness in language models and credit scoring. The most common failures I see involve over-reliance on automation without sufficient controls: missing lineage, weak exception handling, and outcomes that cannot be explained after the fact. In regulated financial environments, those gaps inevitably surface during audits or customer disputes. Institutions that treat traceability, human accountability, and documented change as design requirements, the way mainframe teams always did, will be the ones whose AI systems earn lasting trust.

The most common failures involve over-reliance on automation without sufficient controls. These include missing lineage, lack of exception handling, and the inability to explain outcomes after the fact. In regulated financial environments, these gaps inevitably surface during audits or customer disputes. Regulators are now codifying these expectations, with initiatives like the FCA’s AI lab researching bias in language models and credit scoring. The direction is clear: AI systems in regulated environments will need to demonstrate explainability and auditability in ways that mirror what legacy systems provided by design.

The mainframes I maintained in my early career were not sophisticated by modern standards. But they earned the trust of the organizations that depended on them because every transaction was traceable, every failure was diagnosable, and every change was documented. The agentic AI systems we build today need to meet that same standard. The technology has changed. The requirements for trust have not.


Alexis Peter Francis is a product and data executive with over two decades of experience building enterprise-scale platforms for credit risk, fraud prevention, and regulatory reporting in financial services. He currently leads Customer 360 and Master Data Management platforms that serve hundreds of millions of customers across global markets. His recent research and applied work centers on deploying Agentic AI for automated evidence collection and data validation, with an emphasis on explainability, governance, and sustained human decision-making in highly regulated environments.