Evaluation Framework

How to Evaluate a Governed AI Platform for Healthcare

Six dimensions, forty RFP questions, and the red flags that should end a vendor conversation in the first meeting. Built from evaluations we run inside health system engagements.

Why AI Platform Evaluations Fail

Most AI platform evaluations fail the same way: they score the demo instead of the deployment. Every vendor demos well. The differences that matter show up in month three, when the auditor asks for logs, the CFO asks for usage, and the night shift asks why the tool will not do what ChatGPT did. This framework scores what month three looks like.

We use it as the evaluation layer in our own engagements, and it is vendor-agnostic by design: the right answer is sometimes a tool we do not sell. Where each product category stands today: the independent comparison.

The Six Dimensions

Compliance paper

Dimension 1
  • BAA scope, certifications, and where your data actually goes.
  • Ask: which products and tiers does your BAA cover, exactly? Which subprocessors and model providers touch our data, and do you hold BAAs with each? Provide the chain.
  • SOC 2 Type II or HITRUST report under NDA? The subprocessor question separates real answers from marketing: a platform that routes prompts to model providers must explain how its BAA chain covers that hop.

PHI protection architecture

Dimension 2
  • What happens between the keystroke and the model.
  • Ask: where is PHI detected, masked, or blocked before reaching a model? Is protection default-on or admin-configured? Is any customer data used to train or fine-tune models, under any setting, by you or any subprocessor? Then make them demonstrate it live with a fake patient record.

Multi-model coverage

Dimension 3
  • Whether the platform governs your AI surface or one corner of it.
  • Ask: which model providers are available today? What is your track record for making new frontier models available after release? Can we set which roles reach which models, and disable a provider entirely? Single-vendor tiers answer this dimension by definition; score them honestly against it.

Visibility and audit

Dimension 4
  • The dimension an OCR investigator cares about.
  • Ask: show us the complete log for a single interaction.
  • Can we export audit records to our SIEM? Then the question that ends weak evaluations: produce every AI interaction that touched PHI last quarter, live, during this demo.
  • If the vendor cannot run that query, the dimension is unbuilt.

Adoption reality

Dimension 5
  • Whether staff will actually use it, because governance that gets bypassed governs nothing.
  • Ask: what does time-to-first-value look like for a non-technical clinician? Which usage metrics map to the KPIs the board will ask for? What do your healthcare deployments show at 90 days, with methodology? Any adoption number without a named methodology is a marketing number.

Operations and exit

Dimension 6
  • Who runs it, what breaks, and what leaving costs.
  • Ask: what happens in the first 24 hours after PHI reaches the wrong place? What is the admin burden in hours per week at our size? What do we export if we leave, in what formats? Exit terms are compliance terms: audit records you cannot take with you are audit records you do not have.

Red Flags That End Evaluations Early

"We're HIPAA compliant" with no BAA scope document

  • Compliance is a contract, not a badge.
  • If the paper does not exist, neither does the claim.

No answer to the subprocessor chain question

Your PHI's route matters more than the vendor's front door.

Audit logs described as "on the roadmap"

  • Governance platforms without logs are chat apps with better fonts.
  • Case-study numbers with no methodology belong in the same bin, and so does pricing that only works at 100 percent workforce licensing.

Scoring It

Weight the six dimensions for your organization before any vendor call, because the weighting is the strategy. A 400-bed hospital with a lean IT team weights operations and adoption higher; an academic medical center weights audit and multi-model higher. Score 1 to 5 per dimension against evidence, not demo impressions. Anything scored on a vendor's verbal answer gets a 2 until the document arrives.

Two finalists within a point of each other is not a tie. Rerun dimension 4 with your own auditor in the room and it will not stay a tie.

Evaluating AI Platforms: Common Questions

How long should the evaluation take?

Four to six weeks with the question bank in hand: shortlist from the category map, two-week documentation exchange, demos scored against dimensions 2 and 4 live, references, decision. Evaluations that run past a quarter are usually missing decision rights, not information.

Should we run a formal RFP or a lightweight evaluation?

Under a few hundred seats, the question bank in a structured email beats a formal RFP. Above that, or with board visibility, run the RFP. Either way the questions are the same.

Who should be in the room?

The same seats as your governance committee: security, IT, compliance, clinical, and one frontline user who will live in the tool. Vendors calibrate answers to the most senior compliance person present. Bring them.

What about pilots?

A pilot is dimension 5's evidence, not a substitute for the other five. Time-box it, define its success metrics before it starts, and instrument it with the same KPIs you will run in production.

Get the Full Question Bank

The complete RFP template with all forty questions, scoring sheet, and red-flag checklist is the working document from our health system engagements. Ask and we will send it over, no form maze.