TrustSphere Vendor Assessment: Nuance Gatekeeper, Voice Biometrics and Conversational Fraud Detection for Contact Centres


The telephone is where a bank's fraud controls are weakest and its customers most exposed. Digital channels have device intelligence, behavioural biometrics and a decade of investment. The contact centre has a person, a script, knowledge-based questions long since available on breach dumps, and an agent measured on how quickly the call ends. It is also the channel through which a customer is most likely to be talking to a fraudster while a genuine payment leaves their account.
Nuance Gatekeeper, now part of Microsoft, is the most established answer to that problem. It combines passive and active voice biometrics, watchlists of known fraudster voiceprints, conversational analytics to surface a call going wrong, and an orchestration layer scoring risk across the whole call. It has run at scale in large banks long enough to have a real operational record.
This assessment takes the capability seriously and then asks the harder commercial question. The technology works. Whether it deserves a place on a fraud roadmap now depends on how much of your risk still travels by telephone, how confident you are that a voiceprint means what it meant five years ago, and who will own a fraud control living inside the customer service stack.
Score and Capability Profile
Nuance Gatekeeper scores 6.5 out of 10 in the TrustSphere RiskTech Index, against an index mean of 6.06, in the category of Contact Centre Fraud Prevention and Voice Biometrics.
Taking the capability lines in order, with the standard index lines retained whether or not the vendor addresses them. Voice biometric authentication accuracy scores 9. Known-fraudster voice watchlist and cross-client signal scores 8. Fraud Detection scores 7. Identity Verification and Liveness scores 7. Behavioural Biometrics scores 7, on in-call interaction signal only. Conversational and scam-in-progress detection scores 6. Contact centre and telephony estate integration scores 6. Synthetic speech and deepfake resistance scores 5. Case management and investigator tooling scores 5. Model transparency, bias testing and explainability scores 5. Enterprise Fraud Risk Management scores 4. Device Intelligence scores 3. Transaction Monitoring and Screening scores 3. eKYC and KYB scores 3. Watchlist and Sanctions Screening scores 2.
The fifteen index weights are published so the arithmetic can be checked: voice biometric accuracy leads at 15 percent, and applied to the scores above they give a composite of 6.50.
Read the shape rather than the number: everything tall concerns the human voice inside a live call, everything short the rest of the estate. The risk is not that the product underperforms its own description, but that the description covers less of the fraud problem each year.
That shape also tells a buyer where the integration work will land. A profile this concentrated is not a platform and was never meant to be one, so the composite score should be read as a statement about depth in one channel rather than breadth across a fraud estate. In practice that means the decision is rarely Gatekeeper against another vendor on a shortlist; it is Gatekeeper against the marginal pound spent on payment-time intervention, mule detection on the receiving side or better in-app warnings. Institutions that score it inside a general fraud platform comparison tend to reach the wrong answer in both directions, either dismissing a strong channel control because it lacks case management or buying it in the expectation that it will do work the rest of the estate is failing to do.
What It Actually Does Well
The first genuine strength is removing knowledge-based authentication from the telephone channel without adding friction. Passive voice biometrics verify during the first seconds of natural conversation, so the customer is authenticated before the greeting ends and nobody asks for a memorable word a fraudster has bought. In our review work it was the most reliably defeated control in the contact centre, and those who find it hardest to pass are the elderly and unwell, not the criminal.
The second strength is the known-fraudster watchlist, which separates this vendor from generic voice authentication. Fraudsters are repeat callers by economic necessity, and a voiceprint from an earlier attempt, retained lawfully, identifies the same individual attacking again under a different customer name. Where a consortium view operates across clients the signal strengthens further, and few genuinely cross-institutional fraud signals exist in this channel.
The third strength is the shift from authentication to whole-call risk. The recent positioning is less about verifying the caller at the front door than about scoring the call as it proceeds, combining acoustic conditions, interaction patterns and language signals into a view an agent can act on. Applied to inbound and outbound calls alike, that is a credible attempt to detect an authorised push payment scam while it is happening, which is the only point at which intervention saves money.
Where the Limitations Are
The first limitation should decide the business case: the voice channel is shrinking. Telephone contact continues to fall as a share of interactions in every institution we have reviewed, and the fraud that matters most is increasingly initiated and completed in the mobile app with no human contact at all. The counter-argument deserves weight, because the calls that remain are disproportionately high value, disproportionately made by older and more vulnerable customers, and disproportionately the route by which large irreversible payments are arranged. Covering a small share of interactions is not the same as covering a small share of harm, but that argument must be made with the buyer's own data.
The second limitation is synthetic speech. Voice biometrics were designed for a world in which reproducing a person's voice convincingly took effort, and that world has gone. Cloning from a short sample is trivially available, and the question is no longer whether a voiceprint can be spoofed but how quickly countermeasures adapt as generation tools update. The vendor invests seriously in playback and synthesis detection and is better placed than most, but this is an arms race against an attacker who can iterate cheaply and test offline. A voiceprint should be one factor in a risk decision and never a credential sufficient to authorise a payment.
The third limitation is enrolment coverage, consent and the harm of false rejection. Voiceprints are special category biometric data in the United Kingdom and the European Union, requiring a lawful basis and, usually, explicit consent with a real option to decline. Coverage is never complete, so unenrolled and opted-out customers need a fallback, usually the knowledge-based authentication the programme was meant to retire, and that fallback becomes the attack surface by design. False rejection is not evenly distributed either: illness, distress, speech-affecting conditions, ageing voices and poor lines all raise failure rates, and each describes a customer using the telephone because the app is not available to them. Treating the most vulnerable callers as the most suspicious is both a Consumer Duty problem and a fairness problem.
The fourth limitation combines integration weight, the inferential nature of coached-caller detection, and ownership. Deployment touches session border controllers, the IVR, the recording estate, media streaming, agent desktops and outsourced sites with their own recording law; this is a telephony programme rather than a software installation, and the effort is routinely underestimated. Coached-caller detection is inferential rather than conclusive: the model observes hesitation, stress markers, background acoustics and script-like phrasing and infers coaching, and those signals correlate with manipulation but also with grief, illness and anxiety. Finally, ask who owns it. Gatekeeper lives in the contact centre stack, is usually sponsored by customer experience, and is often justified on handling time savings rather than prevented loss. A fraud control funded from a service budget is exposed whenever that budget is cut.
Questions to Press in the Demo
Do not accept the standard accuracy slide: it uses a controlled population and tells you nothing about your customers. Press on the callers who fail, synthetic speech, the unenrolled, and who runs the system afterwards.
Show us false accept and false reject rates by age band, speech-affecting condition, non-native speaker and line quality, from a live banking client, and what that client did about the worst cohort.
Run a synthetic voice attack on your own enrolled sample in front of us using a current public cloning tool, then give us your countermeasure release cadence and how you learned about the last technique that defeated you.
Describe the fallback path for an unenrolled or opted-out caller, tell us what proportion of calls take it at your largest banking client, and explain why a fraudster would not simply choose it every time.
Take a real coached-caller detection from production, show us the signals the model used, and explain how the bank distinguished manipulation from distress and how many calls it got wrong.
Tell us who tunes this at your reference clients, whether fraud or the contact centre holds the thresholds, and show us the audit record second line would receive.
Verdict
Nuance Gatekeeper is the right purchase for a bank with a large, high-risk telephone channel, an authentication model still resting on knowledge-based questions, and customers who use the phone because they cannot or will not bank digitally. For that institution the 6.5 is fair, and the case rests on removing a control fraudsters defeat routinely.
It is the wrong purchase for a digital-first institution whose voice channel is small and shrinking, for a buyer wanting a broad fraud platform, and for anyone unable to resource a telephony integration programme. It supplies almost nothing outside the call: no payment risk scoring of substance, no device intelligence, no sanctions screening, no identity proofing. It must pair with a payment fraud engine, a digital behavioural and device layer, and case management that treats a voice alert as one input among several. Bought as a contact centre efficiency project that happens to mention fraud, it will drift, and the drift will not show until a claim file needs it.
Suggested Next Steps
Quantify what proportion of your authorised push payment and account takeover losses involve a live telephone interaction, and size the business case on that rather than on call volume.
Require per-cohort false reject testing before contract, define acceptable thresholds for older, unwell and non-native-speaking callers, and review fairness monitoring at the Consumer Duty forum.
Resolve biometric consent, lawful basis and retention, including fraudster voiceprints, with the data protection officer before design, and specify a fallback stronger than knowledge-based questions.
Place ownership and threshold governance with the fraud function, and complete the EBA outsourcing, DORA and operational resilience assessments covering the telephony dependency and exit.
Sources: Financial Conduct Authority, Payment Systems Regulator, Information Commissioner's Office, European Banking Authority, DORA, National Cyber Security Centre, UK Finance, TrustSphere Risk Index, April 2026.
TrustSphere helps financial institutions design and deploy intelligent fraud and financial crime detection solutions. Visit www.trustsphere.ai



Comments