An AI vendor evaluation checklist gives you a sequence for deciding whether an implementation partner can deliver before you commit budget to finding out. This one has seven steps, in the order a real selection happens: define the problem and its current cost, set your target and your walk-away number, treat vendor claims as leads rather than evidence, verify the certifications you can check without technical help, buy a bounded proof on your own data, clear your professional obligations, settle ownership and exit in the contract, and name the person who monitors it after go-live.
Cheap checks come first, so you spend nothing discovering that a vendor will not run a paid proof on your ledger. The expensive commitments come last, once you have evidence you generated yourself.
That emphasis on your own evidence comes out of the research. When academics at Carnegie Mellon scored a corpus of completed vendor disclosure documents in 2026, more than half contained no performance results at all. Vendors described what their systems were for, but rarely described how those systems performed, and none of the evaluations they cited used data anyone else could check. A checklist built on asking vendors better questions runs into that wall.
Why Most AI Vendor Evaluations Fail Before They Start
The standard advice on choosing an AI vendor is a list of questions to ask. The questions are usually fine, but nobody tells you how to read the answers.
Blaine Kuehnert and colleagues at Carnegie Mellon University, working with TechBetter, published the first empirical look at this problem in April 2026. They took 39 completed vendor disclosure documents, each answering the same 23-question template, and scored them against a published transparency rubric. They also interviewed 19 people who write or read those documents, across roughly 20 hours.
The results are worth reading closely if you are about to run a selection process.
On training data and on evaluation, 22 of the 39 disclosures scored the lowest mark available, meaning little or no relevant information. On intended use and context limits, 16 of 39 scored the top mark. Vendors explain what a system is for but go quiet on how it performs.
More than half the documents, 21 of 39, reported no performance results at all. Of those 21, most named a metric and attached no number to it. Where vendors did report figures, they favoured business-value numbers such as time saved, and those figures rarely surfaced any limitation of the system. No evaluation in the corpus used a publicly available dataset, so a buyer holding two of these documents could not compare the vendors or reproduce a single claim.
The interviews explain why, and the explanation is more useful than an accusation would be. Vendors described the disclosure as a place to differentiate themselves, which is what a sales document is for. Detailed disclosure also costs them something real: several named competitive exposure and legal risk, and the smaller firms felt it most. Some vendors could not answer even when willing, because they build on top of foundation models developed by someone else and hold no access to training data, licensing terms, or model internals. Several described themselves as integrators rather than builders, which was accurate.
So the gap is structural, and it does not close because you ask more sharply.
Checklists have their own documented failure mode, and you should know it before you use one. Tom Zick and three co-authors interviewed officials in Brazil, Singapore and Canada who run mature AI procurement regimes, and examined two live procurement checklists in detail. They found that procedures like these drift toward the ceremonial, and the first weakness they name is expertise. A checklist works only when the person running it can judge the answers.
You are a CPA. You are not going to interrogate a model architecture, and a checklist that pretends otherwise wastes your time.
The seven steps below split on that line. The steps needing technical judgement get handed to a structured proof or to an accredited third party. The steps you run yourself are contract terms, professional obligations, and whether anyone wrote down a number before the work started. Those are CPA skills, and they decide more outcomes than the architecture does.
The AI Vendor Evaluation Checklist: 7 Steps

- Write the problem and its current cost down before you contact anyone. One workflow, what it costs your firm today in the numbers your partners argue about, and what better looks like as a figure.
- Set your target and your walk-away number in advance. Decide what counts as success, what you will not spend past, and the date an unproven engagement ends. Do this while you are still calm.
- Treat everything a vendor tells you as the start of the enquiry. Their answers tell you where to look next. Nothing in them is evidence.
- Verify the two things you can check without technical expertise. Certification scope, and who issued the certificate. Ask for the report rather than the badge on the website.
- Buy a bounded proof on your own data, with the pass mark agreed first. Fixed scope, fixed end date, your ledger, and a number you both signed up to beforehand.
- Get ownership and exit in writing, and know which words do the work. A work-for-hire clause may not transfer what you think it transfers.
- Agree who monitors it after go-live, and against what. Name the person inside your firm. The accountability does not travel to the vendor.
Run them in that order. The order is not taken from the air, and the next section explains where it came from.
What the Most Exposed AI Buyers Do Differently
IEEE published a standard for procuring AI and automated decision systems, IEEE 3119-2025. Its scope covers problem definition, solicitation preparation, vendor and solution evaluation, contract negotiation, and contract monitoring, with risk-management methods built for AI rather than borrowed from general IT purchasing. Two of its ideas transfer directly to a mid-market firm. The first is a risk-appetite analysis you complete before evaluation begins. The second is a scoring guide for analysing vendor claims, which exists because the claims need analysing.
The US federal government arrived at a similar shape. OMB Memorandum M-25-22, issued in April 2025, directs agencies to treat AI procurement as a lifecycle, to convene cross-functional evaluation teams, to define performance metrics before soliciting, and to write contract terms that reduce vendor lock-in and preserve the buyer's rights to its own data and to the outputs the system produces.
Neither applies to you. Both were written for high-risk public procurement, and a 92-person firm has no obligation to comply with either. What they give you is the running order. Buyers with the largest budgets, the heaviest legal exposure, and the most public scrutiny landed on the same conclusion, which is that the cheap verifiable checks belong at the front and the commitments belong at the back.
If your Managing Partner asks where the process came from, that is a better answer than a vendor's blog post.
Before the First Call: Define the Problem and the Exit
Step 1. Write the problem and its current cost down before you contact anyone
Pick one workflow and name it. Invoice coding and approval routing. Month-end reconciliation. Document intake from clients. Whichever one bleeds the most hours in your firm right now. What you are avoiding here is the brief that says "AI for the firm," because no vendor can price it and no proof can test it.
Then put three numbers on one page:
- What the workflow costs you today, in units your partner group already argues about. Cost per invoice, hours per close, days from work to bill, the realization gap on a particular service line.
- What it should cost if this works.
- How you will measure the difference, using data you already collect.
The reason this comes first is not tidiness. Whoever defines the metric controls the verdict, and the Carnegie Mellon corpus shows what happens when vendors define it: they report the business-value figure they picked, measured on data they picked, against a baseline they did not disclose. Arrive with your number and you have taken that away from them.
Realization is the honest place to anchor this for most firms. A practice in the $25M fee range sits well below the small-firm benchmarks, and recovering a single point of it is worth six figures at that size. That arithmetic follows from your own standard billings, so run it on your own numbers before you put it in front of the board.
Step 2. Set your target and your walk-away number in advance
Write the exit before you write the brief. Two figures and a date: the spend you will not go past, the pass mark from step 1, and the point at which an unproven engagement stops rather than renews.
Share it with your Managing Partner while nothing is at stake. IEEE 3119-2025 calls the underlying exercise a risk-appetite analysis, and pairs it with a risk register you keep through the engagement.
The reason to do this early has nothing to do with discipline and everything to do with what happens later. Once a pilot is running, the vendor's team is likeable, and your firm has told clients something is coming, the bar moves.
Gartner forecasts that more than 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value, or inadequate risk controls. Read those three as symptoms. Each one is what a missing walk-away number looks like eighteen months in.
Evaluating Vendors: What to Verify and What to Ignore
Step 3. Treat everything a vendor tells you as the start of the enquiry
Your job on a vendor call is not to score answers. It is to work out where to look.
Two questions do most of the work, and both are shaped by what the research found.
Ask which parts of the system the vendor built and which parts they integrate from someone else. A vendor who answers plainly has told you something valuable. A vendor who implies they built the whole stack has told you something too. The Carnegie Mellon interviews found that firms integrating third-party models often cannot answer downstream questions about training data or model behaviour, because the information sits with a provider upstream. That is a limit rather than a failing, and you want to know about it before you write it into a contract.
Then ask for one performance number, with the method behind it. Measured on what data, against what baseline, over what period. Most disclosures in the study named a metric and produced no number. If that is what comes back, write it in your scorecard and move on rather than asking again in a different form.
Watch for the positive signal too, because it is rarer than the red flags and worth more. A vendor who volunteers where their approach breaks, and what they do when it does, is telling you they have shipped something before.
Step 4. Verify the two things you can check without technical expertise
Two documents give you real information, and neither requires you to understand a model.
ISO/IEC 42001:2023 is the first international standard for an AI management system. Certification comes from an independent audit by an accredited certification body, with surveillance audits each year rather than a one-time pass. ANAB in the United States, UKAS in the United Kingdom and RvA in the Netherlands accredit the bodies that issue it, and a further standard, ISO/IEC 42006, sets the requirements those bodies must meet. That last detail is why asking who signed the certificate is a real check rather than a rhetorical one.
For SOC 2, ask which type you are being shown. A Type II report covers whether controls operated effectively across a period. A Type I covers how they were designed at a moment. Ask for the report, read the exceptions section, and treat a logo on a website as nothing.
Now the limit, because overstating this step would mislead you. ISO/IEC 42001 certifies how a vendor manages AI work across its organisation. It does not certify that the system being sold to you performs on your data. The Carnegie Mellon researchers noted that buyers lean on SOC 2 and ISO precisely because no comparable certification exists for model behaviour or fitness for a particular use. That absence is the whole reason step 5 exists.
Step 5. Buy a bounded proof on your own data, with the pass mark agreed first
The Carnegie Mellon authors conclude with a recommendation for buyers: stop expecting documentation to carry the evaluation; pair it with independent testing or a sandbox pilot. For a mid-market firm, that translates into something simpler. Buy a small piece of work before you buy the large one, and set the conditions yourself.
Three conditions, each for a reason.
Bounded. Fixed scope, fixed price, fixed end date. Open-ended pilots are the documented failure mode, and the mechanism is easy to see: with no end date, nobody has to declare a result.
On your own data. A demo runs on the vendor's data, which has been cleaned by people who know what the system struggles with. Your ledger contains your exceptions, your coding habits, and the client who sends 40 invoices as one PDF. Exceptions are where automation breaks, and you cannot find yours in someone else's sandbox.
Pass mark agreed first. The number from step 1, written into the engagement before anyone starts work. Agreeing it afterwards is negotiating with sunk cost on the table.
You get something out of this even when the answer is no. A working artifact, a documented finding you can take to your partner group, and an honest read on the state of your own data, which you will need for any vendor you pick.
Before You Sign: Obligations, Ownership, and Who Watches it
Step 6. Get ownership and exit in writing, and know which words do the work
This is the most concrete item on the list, and the one firms most often get wrong.
Under US copyright law, a work made for hire is either something an employee produces within the scope of employment, or a specially commissioned work that falls into one of nine categories named in the statute, where both parties sign a written agreement saying so. Software is not among the nine categories. A work-for-hire clause on its own therefore may not move copyright in custom code from the vendor who wrote it to the firm that paid for it. A written assignment of copyright does that. The definition sits at 17 U.S.C. §101, and courts have declined to treat commissioned software as work made for hire even where the contract used that exact language.
We are not lawyers, and this is not legal advice. Hand your counsel four questions:
- Does this contract assign copyright to us, or does it only call the work a work made for hire?
- Who owns the outputs the system produces, and who owns the data we put into it?
- What happens to our access if the vendor stops trading? Repository access, escrow, or nothing?
- Can we hire a different firm to maintain this without the original vendor's permission?
That fourth question is the one that decides how much bargaining room you have in year three. M-25-22 tells federal agencies to write terms that reduce lock-in and preserve the buyer's rights to its data and outputs, which is the most demanding buyer in the country treating this as a contract problem rather than a trust problem. A 92-person firm can borrow the posture.
Step 7. Agree who monitors it after go-live, and against what
IEEE 3119-2025 ends its sequence on contract monitoring. The NIST AI Risk Management Framework covers the same ground in its Manage function. Both treat go-live as the middle of the process.
Settle three things before signature. A named person inside your firm who owns the step-1 number. A cadence for measuring it, monthly or quarterly, using the method you defined at the start. And the condition that triggers a review, whether that is a drift in accuracy, a rise in exceptions routed to humans, or a cost that stops falling.
One trap showed up in the Carnegie Mellon interviews and it is worth naming. Several practitioners believed performance monitoring was the vendor's responsibility. Your firm signs the client engagement letter. A vendor can supply the reporting, and a good one will. The obligation stays with you.
AI Vendor Red Flags
Score each vendor on the same rows, in the same room, with the same people present. Comparing notes from three separate conversations you remember differently is how firms end up choosing on rapport.
How Codebridge Works With Mid-Market Firms
We came out of KPMG, which shapes how we scope work. Fixed price and fixed date, agreed before we start, with the deliverable defined in the numbers your board already tracks.
We build on your data rather than demonstrating on ours. The first engagement is a three-week fixed-fee discovery that produces a working prototype against one workflow in your firm, along with an honest assessment of whether the workflow is ready. Some are not, and we would rather tell you that in week two than in month nine.
You own the code. The assignment is written into the contract, not implied by it, which means you can take the work to another firm, maintain it with your own people, or leave it running without us. We have delivered more than 700 projects on that basis.
What we sell is recovered capacity in your practice, measured the way you measure everything else. If that sounds worth fifteen minutes, book a call and we will work out whether one of your workflows fits the three-week prototype. If it does not, we will say so.

Heading 1
Heading 2
Heading 3
Heading 4
Heading 5
Heading 6
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.
Block quote
Ordered list
- Item 1
- Item 2
- Item 3
Unordered list
- Item A
- Item B
- Item C
Bold text
Emphasis
Superscript
Subscript

























