How to test AI so you can trust it
You cannot open the black box, so you test it from the outside. In Part 1 we said there are two ways to make a big AI opinion safe: split it into small steps you control, or benchmark it until you can trust it. This is the half we did not show you. How to test a model against answers you already know are right, how to prove your choices, and what to demand from anyone selling you AI.
As with Part 1, this session runs inside an interactive app rather than a slide deck. You will get a link afterwards so you can click through the demos yourself.
Four sections, one honest premise: if you cannot inspect how a model reasons, you have to judge it by how it behaves. No filler. No vendor spin.
You have never opened an analyst's head either. You train them, review their early files, sample their work and retrain them when they drift. Benchmarking is that same supervision, pointed at a model. Judge the behaviour, not the reasoning.
Your data, your questions, and the answers you already know are right. We walk a real benchmark run and show why there is no best model, only the best model for this task on this data. The winner changes more often than you would expect.
A test proves a point in time. A control keeps proving it. Confidence routing so the model handles what it is good at, audit sampling that stands up, measuring agreement with your own analysts, and knowing when your evidence has gone stale.
The file you need when someone asks how you know it works. Why this model, what you tested, what happens to uncertain cases, how often you re-check, and where you deliberately do not use AI at all.
You will have watched models scored against known-good answers, so you know what to ask for.
Where to draw the line between what AI handles and what a person does, and how to justify it.
The six things you would need to produce if a supervisor asked how you know your AI works.
This session is for people who need to know not just whether AI works, but how to prove it.
Understand what evidence you would need if someone asked how you know your AI works, and the questions to put to any vendor before you sign.
See how confidence routing lets you take the speed where it is safe and keep human judgement where it is not. Automation is a dial you set, not an all or nothing decision.
Watch a real benchmark run, see why the winning model changes with the task, and learn how to build a repeatable evaluation process before you go live.
The foundations: what AI can deliver in KYC today, where black box decisions fail regulators, and how to build auditable workflows.
Request the Part 1 replayWe bring in a regulatory expert to walk the questions a supervisor will put to you about your AI, and the evidence you will need to answer them.
See Part 3