Independent AI evaluation
Independent testing of the AI system before you put it in front of customers.
MosaicAGI is an independent evaluation laboratory for enterprises putting AI into regulated or high-stakes work. We test the system on your own tasks and data, red-team it, and report results to your board whether or not they flatter it.
No account. No card. No sales call.
The problem
Companies are buying AI systems on vendor benchmarks, which are marketing.
Nobody independently tests whether a model actually performs on the buyer's own tasks, how it fails, whether it can be manipulated, or how it behaves on the edge cases that produce regulatory exposure.
Then it goes live, fails publicly, and the board asks who checked.
Who it is for
Chief risk officers
Risk leaders at enterprises deploying AI in financial services, healthcare, insurance, legal and government.
Heads of AI governance
Governance leaders who need a framework for signing off an AI deployment, and the vocabulary to challenge a vendor.
Procurement leads
Buyers of AI systems for regulated or high-stakes functions who need more than vendor benchmarks.
Could you show a board why this AI deployment was signed off?
Answer seventeen questions about a planned or live AI deployment, across performance, manipulation, fairness, the surrounding system and sign-off. See where the evidence is thin and what to ask for first. Runs in your browser. No account, no card, no call.
Method: yes scores a question in full, partly scores half and no scores nothing. Each is weighted from one to three by how much the missing evidence would weaken a sign-off, and the score is the weighted share out of a hundred. Unanswered questions are left out.
Evidence a board or a regulator can read
An independent evaluation laboratory that tests a specific AI deployment on the buyer's own work, and reports the results whether or not they are flattering.
- Built on your own tasks
- An evaluation suite built on your own tasks and data, rather than vendor benchmarks.
- Accuracy with proper statistics
- Task accuracy measured with proper statistics rather than cherry-picked examples.
- Red-teamed for manipulation
- Tested for prompt injection, jailbreaks and data extraction.
- Performance across populations
- Tested for disparate performance across the populations you serve.
- Calibration and the full system
- Calibration and failure behaviour are measured, and the surrounding system is stress-tested, including tool use and retrieval.
- Results reported either way
- A report with pass criteria agreed in advance, and results reported whether or not they are flattering.
Pricing
Start with a free risk assessment. Evaluations are priced by scope, with continuous monitoring available afterwards.
Risk assessment
Freetwo-hour review
For a planned or live AI deployment.
- Failure modes that apply
- Evidence needed for sign-off
Evaluation
RecommendedFrom $35,000to $90,000, depending on scope
Per evaluation engagement.
- Suite built on your own tasks and data
- Red-teaming
- Pass criteria agreed in advance
Monitoring
$12,000per month
Continuous monitoring of a production system.
- Regression testing on production traffic
MosaicAGI never sells or resells AI systems. Prices in USD.
How it works
MosaicAGI, from the first step to the result.
- 01
Start with a risk assessment
A free, structured two-hour review of a planned or live deployment.
- 02
Agree the pass criteria
In advance of testing.
- 03
Evaluate the deployment
Accuracy, red-teaming, disparate performance, calibration and the surrounding system.
- 04
Receive the report
Written for a regulator or a board, with results reported whether or not they are flattering.
Questions people actually ask
- What is the free Deployment Risk Assessment?
- A structured two-hour review of a planned or live AI deployment. It identifies the specific failure modes that apply, what evidence would be needed to sign it off, and which of the current evaluation claims are actually meaningless.
- Why not rely on vendor benchmarks?
- Vendor benchmarks are marketing. An evaluation built on your own tasks and data tests how the system performs, how it fails and whether it can be manipulated.
- What does an evaluation test?
- Task accuracy with proper statistics, prompt injection, jailbreaks and data extraction, disparate performance across the populations you serve, calibration and failure behaviour, and the surrounding system including tool use and retrieval.
- What if the results are unflattering?
- They are reported anyway. Pass criteria are agreed in advance, and results are reported whether or not they are flattering.
- What does an evaluation engagement cost?
- 35,000 to 90,000 dollars per engagement, depending on scope.
- Can a system be monitored after it goes live?
- Yes. Continuous monitoring with regression testing on production traffic is 12,000 dollars a month.
- Does MosaicAGI sell AI systems?
- No. It never sells or resells AI systems, which is what preserves its independence.
- Who is it for?
- Enterprises deploying AI in regulated or high-stakes functions, including financial services, healthcare, insurance, legal and government.
- Why should I trust the AI deployment evidence check?
- The score comes from your answers, worked out in your browser. It does not test the system; it shows which evidence exists. The free Deployment Risk Assessment reviews the deployment itself and names the failure modes that apply to it.
Could you show a board why this AI deployment was signed off?
Answer seventeen questions about a planned or live AI deployment, across performance, manipulation, fairness, the surrounding system and sign-off. See where the evidence is thin and what to ask for first. Runs in your browser. No account, no card, no call.
Open the free toolIt runs in your browser. MosaicAGI never sees your inputs.