Not every step should be agentic in KYC

By Henry Bruce | 2 hours ago
Assess which decisions a KYC agent gets to make

In our API vs MCP blog we explored why banks need both application programming interface (API) and Model Context Protocol (MCP) services. We also explored what can go wrong when a KYC agent chooses its own tool calls.

But how do banks assess these decisions for their KYC agent readiness to capitalize on the AI opportunity?

Identify the decisions in your workflows

Maybe this sounds familiar. Business and technology teams are carefully designing a future-state client onboarding workflow. In it, KYC agents make many of the decisions. The question in front of them is how much of that future state they can build today. They point to the target entity identification step. Could an agent make its decisions?

In the workflow it reads as one step. A client gives you a name and some identifying detail, and you establish which real company it refers to. But look more closely and this step is composed of a number of decisions:

  • How to search, and on which of the supplied details.
  • Which registries, exchanges and regulators are in scope for the jurisdictions implied.
  • Whether a returned candidate is close enough to be worth considering.
  • Which of the plausible candidates is the target.
  • Whether the evidence for that choice is strong enough to proceed.

One decision may be ready to be made agentically now and another not for a long time. The teams at this bank go back through the design and annotate each step with the decisions inside it. How should they then assess the readiness of each to be made agentically?

Four questions to ask of each decision before handing it to a KYC agent

Ask four questions of every decision. And if a decision fails one or more questions, you have its backlog.

  • Bounded: Are the agent’s choices clear? Two aspects of the choices must be defined: the options available to the agent, and the basis on which it chooses among them. A defined menu with an undefined selection rule is not a bounded choice, because the agent is still deciding on a basis nobody wrote down.
  • Observable: Can we see which option the agent chose, and on what basis? The decision record needs to carry the option the agent settled on and the evidence behind it: the alternatives it was choosing among, and what distinguished the one it took.
  • Detectable: Would we know if the agent made the wrong choice? Something must indicate that the choice was wrong without anyone looking, and it has to do so before the decision is relied on. “We would find out eventually” does not qualify, because the test is whether you find out before the decision is used, not whether you find out at all.
  • Recoverable: Can an agent’s wrong choice be made right before it matters? The choice can be undone and the work that followed it redone without adversely affecting the outcome.

Why detectability is the question most often missed

The design guidance on agent bounding and observability is well established: an agent’s permitted actions and reachable resources should be limited, and its actions should be observable, with planning and reasoning recorded. There’s guidance on recoverability (sometimes called reversibility) too, but this is generally only framed as a trigger for human confirmation of a critical action rather than as a test of whether a decision should be made agentically.

In contrast, detectability is largely absent from the guidance and not without reason. In many agent deployments to date, a wrong answer comes to light through use. A software unit test fails. A wrong booking produces a complaint. A bad fraud decision produces a chargeback. In KYC, the only parties who know the entity was wrong are the client, who never sees the file, and a criminal, who has no reason to say. The signal does arrive in the end, through an enforcement action or a question from a correspondent bank, but that could be months or years later. All four questions are important to ask together, but detectability stands out as the one least likely to have already been thought about.

Assessing the plausible candidate’s decision

Back to the fourth decision in that workflow: which of the plausible candidates is the target?

  • Bounded: yes, provided the candidates come from a deterministic search rather than the agent’s own recall, so the option set is defined. The second aspect is where most banks find work to do, particularly in the edge cases. Ensure that none of the match criteria, the thresholds, the identifiers treated as authoritative, or any allow or deny lists are being left to the agent’s discretion – a defined candidate list with some undefined selection rules doesn’t qualify.
  • Observable: yes, provided the record carries the candidate set and the evidence, and not only the entity the agent settled on.
  • Detectable: no. Imagine the agent picks candidate B. The truth is candidate A. Both are real companies with similar names. The ownership structure of B is retrieved and comes back complete and valid. Its beneficial owners are identified, correctly. Screening runs on those beneficial owners and returns clean. Every call succeeds, every fact carries its source and retrieval date, and the audit trail is accurate – it records exactly what happened. Everything the agent retrieves after the identification inherits the identification, so nothing retrieved later can contradict it and the error is self-confirming. Contradicting evidence might exist in the file, but nothing compares the agent’s conclusion with it. When the error is found, remediation would reach every file identified the same way.
  • Recoverable: yes, but it doesn’t help.An identification is trivially reversible, and reversibility is worth nothing on an error nobody knows about. (Nor does explainability close the gap: the explanation is accurate and the choice it explains is still wrong.)

Three of the four hold. This decision isn’t blocked from going to a KYC agent but detectability needs work.

Adding a second, independent assertion as a detectability signal

The instinct at this point, and the one those teams will hear first, is to make the agent better at identification. That won’t help. A better agent gets the answer right more often; it still can’t tell you when it got the answer wrong. Detectability is a property of the system around the agent rather than of the agent itself, which means the work is to add a signal.

The step produces exactly one assertion about which company this is, and everything after it inherits that assertion without testing it. So, the addition is a second, independent assertion that must agree, and a defined response (which could be escalation) when it doesn’t. What’s missing is the comparison signal. Which comparison is right depends on what the bank already holds.

One option is to identify twice, on independent evidence: once on name and address, separately on a supplied identifier. Both paths must arrive at the same entity, and disagreement is the signal. That’s ordinary practice in record linkage; what’s less usual is treating the disagreement as detection rather than as an input to better matching.

Another option is to flag the condition rather than the error. Where more than one candidate scores above the threshold, that’s a detectable condition even when you can’t say which choice is wrong. The discipline is to work backwards from the error rates rather than forwards from the threshold: decide first how often you are prepared to link two different companies, and how often you are prepared to miss a genuine match and let those two numbers set the thresholds. Anything falling in between goes to a person.

The trade-off: more cases for review

The cost is review volume. Independent identifications will sometimes disagree, and those disagreements get escalated. Tighter error rates catch more errors but send more cases to review. Sometimes no second assertion exists at all. Take several similarly named entities in one jurisdiction, and a client who supplied a name and a country. Here, there may be no other The fix is to change how the data is collected and processed in the first place.

Knowing which decisions need a signal

Comparing two independently derived answers and investigating the difference is reconciliation. It has been standard practice in banking for decades. A decision that’s bounded, observable and recoverable but not detectable needs reconciliation as a signal, not a control. Without the four questions there’s nothing telling a team which of their annotated decisions those are.

Which is what the teams in our workflow can now do with their annotations: take each decision, put the four questions to it, and see what falls out.

Question What it asks What a no means
Bounded Are the agent’s options, and its basis for choosing between them, both defined outside the agent? Write down the selection rules, rather than just the option list
Observable Does the record show which option the agent took, and on what evidence? Return the alternatives and the evidence, and not only the answer
Detectable Would something show the choice was wrong before it is relied on? Add a second, independent assertion that has to agree; set thresholds by modeling error rates
Recoverable Can the choice be undone, and the work that followed it redone, in time to matter? Require human confirmation before the action that can’t be undone

 

The ones that fail only on detectability aren’t completely blocked; they likely need an extra signal.

That’s why EC360’s MCP services return the candidate set and the full provenance with evidence behind the selection. A second assertion must have something to disagree with, and a bare answer doesn’t give an agent anything.

________________________________________
References
• Santiago Díaz, Christoph Kern and Kara Olive, Google’s Approach for Secure AI Agents: An Introduction, Google, May 2025 https://storage.googleapis.com/gweb-research2023-media/pubtools/1018686.pdf
• OpenAI, A practical guide to building agents https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf
• Anthropic, Building effective agents, December 2024 Building Effective AI Agents

 
Author: Henry Bruce

Henry Bruce is Principal AI Engineer at Encompass, working at the intersection of data and AI to shape the next generation of KYC experiences. His work spans the architecture and delivery of agentic KYC, including the EC360 MCP services, which connect customer AI environments and their agents to trusted, AI-ready corporate digital identity data. Before Encompass, Henry was the founding product engineer at a climate tech startup, building AI-native decision intelligence grounded in complex data.

You also might be interested in

west
east

Discover corporate digital identity from Encompass

 

Find out more