Compliance
700,000 calls a month, 120 deals a day: why we built ComplAI
One deployment now runs 700,000 calls a month and closes around 120 deals a day, each one owing compliance to a regulator, a supplier and an industry code. No team can check that. This is what we built instead.

- Author
- Team xAIa
- Published
- 25 July 2026
- Reading time
- 9 min read
Last month, one deployment of ours carried 700,000 conversations. On its heaviest days it runs 50,000 calls between nine in the morning and six in the evening. Out of that volume, a UK energy brokerage closes somewhere around 120 deals a day, every one of which is a regulated transaction that somebody, eventually, may be asked to justify.
ComplAI is xAIa's proprietary compliance intelligence platform. It reviews every recorded call against the regulations, scripts and supplier rules that govern it, cross-checks each verdict across several AI agents, and turns the result into an auditable record you can defend. This is the problem that made us build it.
Their compliance function is a handful of people. Working flat out against a scorecard, a reviewer clears fifteen to twenty calls a day.
Put those two numbers next to each other and the situation stops being a resourcing question. There is no headcount that closes that gap, at any budget, in any market. This is not a failure of a compliance team, and we have now seen the same ratio in insurance, in collections, in banking and in clinics. The volume of regulated conversation a modern operation produces passed the volume a human function can inspect some time ago, and most businesses have quietly agreed not to look directly at it.
We had also made it worse ourselves. Our voice AI answered everything, worked through the night and never tired, so volume climbed while the inspectable share of it collapsed. We had solved the conversation and created a supervision problem underneath it.
The unit of compliance is not the call
The first thing we got wrong was the shape of the problem. We assumed a compliance system reviews calls, because that is what QA software does.
A deal does not live in one call. It opens on one conversation and closes on another, frequently with a different agent, sometimes weeks apart, moving through phases that differ depending on which supplier it lands with. The obligation attaches to the deal, not to any single recording inside it. A disclosure made on the opener and contradicted on the closer is a breach that no per-call review will ever catch, because both calls look clean in isolation.
So the system had to reconstruct the deal first and assess the conversation as a whole: every call attached to it, in sequence, across suppliers, through each state the deal passed on its way to being signed. Around 120 of those assembled a day, each one a small evidentiary case file.
Three rule sets, all in force at once
The second thing that made this harder than it looks is that a deal does not owe compliance to one authority.
It owes compliance to the regulator, whose rules govern what must be said, when, and in what terms. It owes compliance to the supplier, because each one carries its own contractual requirements about how its products may be represented, and those requirements differ from supplier to supplier and change without much notice. And it owes compliance to the industry code and the wider body of regional law sitting above both.
Those three layers do not agree with each other neatly. A phrase that satisfies the regulator can breach a supplier's terms. Encoded across a live deployment, that adds up to hundreds of documents, and the obligations inside them are not paraphrasable. Compliance turns on specific words in specific positions, so a system that reads a summary of the rules is checking against a rumour of the rules.
Opinion, or record
Somebody on a call asked the obvious question early on. If a language model can hold a compliant conversation, why can it not read one?
The honest answer then was that it could, badly. You can hand a transcript to a model and ask whether the call was compliant, and it will tell you. It will also tell you something slightly different an hour later, cite nothing you can verify, and produce a judgement with no standing whatsoever in front of a regulator. That is not a compliance function, it is a machine having an opinion about your business.
The distance between an opinion and a record turned out to be the entire product, and closing it took two things: a brain, and more than one reader.
The brain
The first thing we built was not a model. It was a way of writing obligations down so a machine could act on them precisely: standards, rejection codes, severity levels, risk tags and call types, each one traceable back to the exact clause of the exact document it came from, the whole thing version-controlled so you can always prove which rules were in force on the day a call was made.
Hundreds of source documents go in, read at the level of the individual word rather than compressed into a summary. That first deployment settled at eight standards and twenty-seven rejection codes sitting on top of that corpus. Building it took longer than building anything else in the platform, it has to be rebuilt for every industry we enter, and it is the reason ComplAI produces findings a compliance officer can defend rather than sentiment about a phone call.
More than one reader
The verdict does not come from a single model reading a transcript, because a single model reading a transcript is exactly the thing that produces confident nonsense.
Multiple agents examine the same reconstructed deal from different positions. One holds the regulatory obligations. One holds the supplier's specific contractual requirements. One holds the industry code and regional law. Each returns its own findings against its own layer, and those findings are then cross-checked against each other before anything reaches a person.
The interesting part is what happens when they disagree. A call one agent passes and another fails is not a problem to be averaged away, it is the most informative output the system produces, and it goes to a human with both positions attached. Some of the sharpest compliance issues we have surfaced came out of exactly that disagreement, in cases where a call satisfied the regulator cleanly and quietly breached a supplier's terms.
There is one more piece of machinery worth naming. An obligation has a place in a conversation, not just a presence in it. A mandatory disclosure read at the end of a call is not equivalent to one read before the customer committed to anything, and a system that searches the transcript for the phrase will pass a call that failed. Every conversation is broken into its structural phases first, and each obligation is tested where it was supposed to occur.
The part we got wrong first
We assumed the value sat in the detection, and that better flagging was the whole job.
Reviewers do not trust a verdict they cannot verify in five seconds. Early on, a finding would say a disclosure was missing at 11:42, and the reviewer would open a forty-minute recording, scrub around, lose their place, and end up listening to the whole thing anyway. The system was right and it was saving nobody any time.
So the first feature we shipped, ahead of multi-tenancy and ahead of the dashboards, was an audio cutter. A reviewer jumps to the moment, hears the eight seconds under dispute, and exports only that segment into a case file instead of handing a colleague a forty-minute recording and a timestamp. Adoption changed the week it went in. Trust in an AI verdict comes from verifiability far more than from accuracy, and we would not have guessed the order of those two.
What changed on the floor
Manual compliance effort on that deployment fell by more than 80%. Reviewers stopped hunting and started confirming, opening a queue already assembled by deal, ranked by risk, with the evidence attached and the moment marked. The hours that came back went into the borderline cases, which were always the ones worth arguing about.
Every recording carries its own audit panel: who uploaded it, who it was assigned to, who reviewed it, who listened to it, and when. Supervisors got something they had never had, which is performance data on the review function itself, so compliance capacity became a number to manage rather than a feeling to worry about.
A compliance report tells you what a sample sounded like. A record tells you what every conversation contained. Only one of those survives the question.
The change that matters most stays invisible until the day it is needed. When a customer disputes a sale, the deal, its calls, its verdicts and the rules in force that week are already assembled. When a regulation or a supplier's terms change, the archive is re-scored against the new rule and the exposure is known before anyone outside the business asks. Where that evidence physically sits matters too, which is a separate argument we have made at length.
Who audits the AI
This closes the loop we opened at the start. The AI employees placing those 700,000 calls a month are now reviewed on the same terms as the humans, against the same brain, producing the same evidence.
Very few vendors deploying AI agents will audit their own output, and it is a fair question for any board to put to one before signing. It sits alongside the governance layer we deploy through i-GENTIC, which controls what an agent is permitted to do at the moment it acts. One prevents, the other proves, and regulated operations need both.
Where it goes from here
The engine does not care what it is pointed at. What changes is the brain: Ofgem obligations and supplier terms for an energy brokerage, CBUAE conduct and telemarketing rules for a bank or an insurer, DHA requirements for a clinic, TDRA rules and the Do Not Call Registry for anyone running outbound in the UAE.
It does not have to stay on the phone either. Anything that puts an obligation next to a record can be assessed the same way: a submitted document, a chat thread, a case file, an application. Voice at 50,000 calls a day was the hardest place to start, which is the reason we started there.
We built ComplAI because we needed it before anybody asked us for it. That is usually the only good reason to build anything.
Want to see ComplAI run against a batch of your real calls and your own rule set? Speak to us and our AI will call you back within a minute. That's the demo.




