Skip to content

Independent AI evaluation

Independent testing of the AI system before you put it in front of customers.

MosaicAGI is an independent evaluation laboratory for enterprises putting AI into regulated or high-stakes work. We test the system on your own tasks and data, red-team it, and report results to your board whether or not they flatter it.

No account. No card. No sales call.

The problem

Companies are buying AI systems on vendor benchmarks, which are marketing.

Nobody independently tests whether a model actually performs on the buyer's own tasks, how it fails, whether it can be manipulated, or how it behaves on the edge cases that produce regulatory exposure.

Then it goes live, fails publicly, and the board asks who checked.

Who it is for

  • Chief risk officers

    Risk leaders at enterprises deploying AI in financial services, healthcare, insurance, legal and government.

  • Heads of AI governance

    Governance leaders who need a framework for signing off an AI deployment, and the vocabulary to challenge a vendor.

  • Procurement leads

    Buyers of AI systems for regulated or high-stakes functions who need more than vendor benchmarks.

Could you show a board why this AI deployment was signed off?

Answer seventeen questions about a planned or live AI deployment, across performance, manipulation, fairness, the surrounding system and sign-off. See where the evidence is thin and what to ask for first. Runs in your browser. No account, no card, no call.

The deployment

Answer for one planned or live deployment, as things stand today. Nothing is sent.

Performance on your own tasks

Has the system been tested on your own tasks and data, not only on vendor benchmarks?

Are accuracy results reported with proper statistics, such as sample size and uncertainty?

Were pass criteria agreed before testing began?

Do you know how the system fails, not only how often it is right?

Manipulation and data

Has it been tested for prompt injection, including through documents and retrieved content?

Has it been tested for jailbreaks against the behaviour your policies forbid?

Has it been tested for data extraction?

Fairness and calibration

Has performance been compared across the populations you serve?

When the system says it is confident, is it right that often?

Is it defined and tested what the system does when it is unsure?

The surrounding system

Have the actions the system can take through tools been tested, not only its answers?

Has retrieval been tested, including when it returns the wrong document or nothing?

Will performance be monitored after go-live, with regression tests on production traffic?

Sign-off

Is a named person accountable for signing off the deployment?

Could a regulator or board read the evidence and follow it?

Was the evidence produced independently of the vendor?

Have you checked which of the vendor's evaluation claims apply to your own use?

Answer the questions to see your score. It updates as you go.

Method: yes scores a question in full, partly scores half and no scores nothing. Each is weighted from one to three by how much the missing evidence would weaken a sign-off, and the score is the weighted share out of a hundred. Unanswered questions are left out.

Evidence a board or a regulator can read

An independent evaluation laboratory that tests a specific AI deployment on the buyer's own work, and reports the results whether or not they are flattering.

Built on your own tasks
An evaluation suite built on your own tasks and data, rather than vendor benchmarks.
Accuracy with proper statistics
Task accuracy measured with proper statistics rather than cherry-picked examples.
Red-teamed for manipulation
Tested for prompt injection, jailbreaks and data extraction.
Performance across populations
Tested for disparate performance across the populations you serve.
Calibration and the full system
Calibration and failure behaviour are measured, and the surrounding system is stress-tested, including tool use and retrieval.
Results reported either way
A report with pass criteria agreed in advance, and results reported whether or not they are flattering.

Pricing

Start with a free risk assessment. Evaluations are priced by scope, with continuous monitoring available afterwards.

  • Risk assessment

    Freetwo-hour review

    For a planned or live AI deployment.

    • Failure modes that apply
    • Evidence needed for sign-off
    Get in touch
  • Evaluation

    Recommended

    From $35,000to $90,000, depending on scope

    Per evaluation engagement.

    • Suite built on your own tasks and data
    • Red-teaming
    • Pass criteria agreed in advance
    Talk to us
  • Monitoring

    $12,000per month

    Continuous monitoring of a production system.

    • Regression testing on production traffic
    Talk to us

MosaicAGI never sells or resells AI systems. Prices in USD.

How it works

MosaicAGI, from the first step to the result.

  1. 01

    Start with a risk assessment

    A free, structured two-hour review of a planned or live deployment.

  2. 02

    Agree the pass criteria

    In advance of testing.

  3. 03

    Evaluate the deployment

    Accuracy, red-teaming, disparate performance, calibration and the surrounding system.

  4. 04

    Receive the report

    Written for a regulator or a board, with results reported whether or not they are flattering.

Questions people actually ask

What is the free Deployment Risk Assessment?
A structured two-hour review of a planned or live AI deployment. It identifies the specific failure modes that apply, what evidence would be needed to sign it off, and which of the current evaluation claims are actually meaningless.
Why not rely on vendor benchmarks?
Vendor benchmarks are marketing. An evaluation built on your own tasks and data tests how the system performs, how it fails and whether it can be manipulated.
What does an evaluation test?
Task accuracy with proper statistics, prompt injection, jailbreaks and data extraction, disparate performance across the populations you serve, calibration and failure behaviour, and the surrounding system including tool use and retrieval.
What if the results are unflattering?
They are reported anyway. Pass criteria are agreed in advance, and results are reported whether or not they are flattering.
What does an evaluation engagement cost?
35,000 to 90,000 dollars per engagement, depending on scope.
Can a system be monitored after it goes live?
Yes. Continuous monitoring with regression testing on production traffic is 12,000 dollars a month.
Does MosaicAGI sell AI systems?
No. It never sells or resells AI systems, which is what preserves its independence.
Who is it for?
Enterprises deploying AI in regulated or high-stakes functions, including financial services, healthcare, insurance, legal and government.
Why should I trust the AI deployment evidence check?
The score comes from your answers, worked out in your browser. It does not test the system; it shows which evidence exists. The free Deployment Risk Assessment reviews the deployment itself and names the failure modes that apply to it.

Could you show a board why this AI deployment was signed off?

Answer seventeen questions about a planned or live AI deployment, across performance, manipulation, fairness, the surrounding system and sign-off. See where the evidence is thin and what to ask for first. Runs in your browser. No account, no card, no call.

Open the free tool

It runs in your browser. MosaicAGI never sees your inputs.