Running the Evaluation Stage Without Being Sold To: Scoring, Proof of Concept and the Demonstration in 2026


A scripted demonstration on the vendor's own data proves exactly one thing: that the vendor is capable of preparing a demonstration. It proves nothing about detection quality on your portfolio, nothing about the effort your engineering team will spend on integration, nothing about what the platform costs in its second year, and nothing about whether the people who will actually operate it can do so without a consultant sitting behind them. Yet the scripted demonstration remains the single most heavily weighted input into technology selection decisions across financial crime and fraud, and it is the input with the lowest evidential value of any that is routinely collected.
The purpose of the evaluation stage is to replace assertions with evidence. Almost every common practice does the opposite. Scoring models are built after responses arrive, which allows weightings to be shaped around a preferred answer. Proofs of concept are run on vendor infrastructure with vendor data and measured by the vendor. Reference calls are conducted with references the vendor selected, who have been briefed, and who are asked questions general enough to be answered pleasantly. Total cost is assessed on licence fee alone, which excludes the majority of what the institution will actually spend.
None of this is dishonest on the vendor's part. A vendor's job is to present its product in the best available light, and a well-run sales process will do that skilfully and legitimately. The failure is on the buying side, and it is a failure of process design rather than of judgement. What follows is how to design the evaluation stage so that the evidence it produces is worth something, drawing on TrustSphere's experience of running and observing selection processes in fraud and financial crime technology.
Where Procurement Goes Wrong
The first and most common failure is the retrofitted scoring model. Requirements are issued, responses arrive, the panel reads them, forms a view, and only then agrees the weightings. At that point the weightings are no longer a statement of what the institution needs; they are a mechanism for producing the preferred result. This is rarely deliberate and it is almost always visible afterwards in the audit trail, which is a problem in its own right where the procurement is governed under outsourcing and operational resilience requirements. The tell is a weighting scheme with unusual precision in exactly the areas where the front-runner is strong.
The second failure is treating the demonstration as evidence. A demonstration is a rehearsed performance on a curated dataset, and every vendor in a competitive process will present one that works. The variance the panel observes between demonstrations is largely variance in presentation quality, in the seniority of the pre-sales engineer, and in how much preparation time each vendor invested. None of these correlate reliably with product quality, and two of them correlate inversely with the vendor's likely attentiveness after contract signature.
The third failure is the proof of concept that proves nothing. This occurs when success criteria are not agreed before the exercise starts, when the data used is synthetic or vendor-supplied, when the vendor runs the environment and produces the results, and when the exercise concludes with a presentation of findings rather than a measurement against a pre-agreed standard. A proof of concept designed this way cannot fail, which means it cannot inform a decision. TrustSphere's observation across selection processes is that the majority of proofs of concept in this market fall into this category, and that the institutions running them frequently describe them as rigorous.
The fourth failure is scope. Evaluations concentrate on functional capability, which is the part vendors compete on, and skip implementation reality, which is where most of the cost and nearly all of the disappointment lives. The questions that predict whether a programme will succeed are unglamorous and are asked too late: who performs the integration, what does the institution have to build itself, what happens in year two, and what does it cost to change something once the vendor's implementation team has gone.
Running the Stage Properly
Fix the scoring model before responses arrive, and put it in writing. Weightings should be derived from the requirements and agreed by the panel and the accountable executive before the request for proposal is issued, not after. Each criterion needs a defined scoring scale with described anchor points, so that a score of three means something specific rather than a general feeling. Where a weighting is changed after responses arrive, the change and its rationale should be documented and approved, which raises the cost of doing it for the wrong reason without prohibiting it for the right one.
Design the proof of concept on your own data, with success criteria agreed in writing before it starts. That means a defined dataset drawn from your own historical population, defined measures (detection of known outcomes, novel detections assessed by your investigators, false positive rate at a stated operating threshold, latency at your peak volume), a defined threshold for what constitutes a pass, and a defined duration. Agree in advance what happens if the vendor asks for more tuning time, because they will, and the answer should be recorded before anyone has an interest in it.
The institution, not the vendor, must measure the result. The vendor may operate the technology, but the dataset should be controlled by the institution, the outcome labels should be withheld, and the analysis should be performed by the institution's own analytics function or an independent party. A vendor measuring its own proof of concept will produce a result that is technically accurate and selectively framed. This is not cynicism; it is the same principle that stops a firm from letting a supplier audit itself, and it is well understood everywhere in the institution except technology selection.
Ask the implementation questions early, in writing, and score them. Who performs the integration and with whose staff. What must the institution build itself, specifically, listed as deliverables. What is the data engineering effort and who has assessed it. What are the licence, support, professional services, infrastructure and change costs in years one, two and three, and what triggers a price increase. What does it cost to add a new business line, a new country, or a new data source after go-live. How many of the vendor's staff assigned to your implementation will still be on the account after six months. These are difficult to answer with a slide, which is why they belong in the written response where the answers are attributable.
Run reference calls that can produce an honest answer. Vendor-supplied references are worth having but should never be the only ones; ask for references matched to your size, region and use case, and separately identify your own through professional networks and industry forums. Speak to the operational owner rather than the executive sponsor, ask questions that assume problems occurred ("what surprised you in month three", "what did you have to build that you had not expected", "what would you do differently"), and ask what the vendor was like when something went wrong. A reference who has nothing critical to say has either had an unusually good experience or is not speaking freely, and it is worth establishing which.
Evaluating What You Get Back
Read responses for evidence rather than for capability claims. A capable response distinguishes between what the product does today in production at named client scale, what is configurable, what is on the roadmap, and what would be custom development. Vendors that blur those categories should be pressed to separate them explicitly and in writing, and the panel should score the answer that comes back rather than the original response. A roadmap commitment is not a capability, and a capability demonstrated once at one client is not the same as a standard deployment.
Score total cost of ownership including internal effort, because internal effort is usually the largest single line and is almost never in the comparison. Build the model to include institution-side engineering and data work, the investigative or operational headcount implied by the alert volume the product will generate, model validation and governance effort, ongoing tuning, and the internal cost of the implementation programme itself. Two products with similar licence fees can differ by a factor of several in total cost once these are included, and the cheaper licence is frequently the more expensive product. Under DORA and the EBA outsourcing guidelines the institution also needs to evidence exit and substitutability planning, so the cost of leaving belongs in the model as well.
Keep the panel honest when there is an internal favourite, because there usually is, and the favourite is often the right answer. The problem is not preference, it is unexamined preference. Practical measures work better than exhortation: have panel members score independently before any group discussion, record scores before the discussion begins, require written justification for any score more than one point from the panel median, and appoint someone explicitly to argue the case against the leading candidate. Where a panel member has a prior relationship with a vendor, record it and let them score with it declared rather than pretending it does not exist. The objective is not to eliminate judgement, which would be undesirable as well as impossible; it is to make judgement visible and therefore challengeable.
Conclusion
The evaluation stage is the only point in the procurement where an institution has leverage, information asymmetry in its favour, and a genuine ability to test claims. It is also the stage most often run on autopilot, with a scoring model assembled to fit a conclusion, a proof of concept designed so that it cannot fail, and references chosen by the party being evaluated. Every one of those choices makes the process more comfortable and the decision worse.
Running it properly is not expensive and it is not adversarial. It requires deciding what matters before you know who will provide it, testing on your own data with your own measurement, asking implementation questions while the vendor still has an incentive to answer them accurately, and building enough structure into the panel that preference has to be argued rather than assumed. Vendors who are confident in their product tend to welcome this, because it distinguishes them from competitors who are confident in their demonstration. The ones who resist it have told you something useful for free.
Suggested Next Steps
Agree and document the scoring model, weightings and anchored scoring scales with the accountable executive before the request for proposal is issued, and require documented approval for any weighting change made after responses arrive.
Design proofs of concept on institution-controlled data with withheld outcome labels, written pass criteria, a fixed duration, and analysis performed by your own analytics function or an independent party rather than the vendor.
Add a scored implementation and commercial section to the written response covering integration ownership, institution-side build deliverables, data engineering effort, year two and year three costs, price increase triggers, and the cost of adding a business line or data source after go-live.
Build a total cost of ownership model that includes internal engineering, operational headcount implied by alert volume, model validation and tuning, and exit and substitutability costs required under DORA and outsourcing guidelines, and score independently before group discussion with an assigned challenger to the leading candidate.
Sources: European Banking Authority guidelines on outsourcing arrangements; Digital Operational Resilience Act requirements on ICT third-party risk, contractual provisions and exit strategies; Financial Conduct Authority and Prudential Regulation Authority operational resilience and outsourcing policy; Financial Conduct Authority and Prudential Regulation Authority model risk management expectations; Wolfsberg Group guidance on effective monitoring and technology governance; TrustSphere Risk Index, April 2026.
TrustSphere helps financial institutions design and deploy intelligent fraud and financial crime detection solutions. Visit www.trustsphere.ai



Comments