AI Evaluation Baseline

Know whether your AI actually works, on your data, at your accuracy bar, before it touches production.

What this engagement does

We build an evaluation dataset from your historical records and test candidate AI configurations against it. You receive a report explaining what passes, what fails and why. We agree acceptance criteria with your subject-matter experts before evaluation begins.

What you receive

  • An evaluation dataset built from your historical data. Your subject-matter experts curate it so "correct" reflects your business rules.
  • A repeatable, automated evaluation harness you own, runnable on every model update, prompt change, or vendor claim
  • Evaluation report with failure taxonomy: why each failure happened and what that implies for deployment scope
  • Compare measured cost per run across frontier and smaller models.
  • Go/no-go recommendation per workflow: automate, automate with human approval, or leave it deterministic
  • Your team receives instructions for running the harness, adding cases and interpreting results.

How the work progresses

  1. Scope and access

    Walk through the workflow with your subject-matter experts and arrange read-only access. Agree accuracy and cost limits in writing.

  2. Build the evidence

    We curate and label historical records with your subject-matter experts, using your business rules to define correct results.

  3. Review and validate

    Build an automated evaluation suite that runs candidate models against the evaluation dataset and measures cost per run.

  4. Handover

    Review the findings with your decision makers in a working session. Receive the report, runbook, harness and all evaluation artifacts.

Working together

  • We agree scope and price, arrange read-only access, and define acceptance criteria with your subject-matter experts before evaluation.
  • Senior architect does the work directly, no handoff
  • Data stays in your environment where required. We agree healthcare data handling requirements with your team.
  • Built on standard, vendor-supported tools that your team can maintain. You own the deliverables and can operate them without Lotus.

What can follow

  • The Agent Deployment Sprint takes one workflow into production using your evaluation baseline as the acceptance test.
  • Recurring evaluation runs on model or vendor changes, priced per run
Discuss this engagement

What needs to
work better?

Tell us about the system, the people who use it, and what is getting in their way. We’ll help define a useful first step, with a clear scope and acceptance criteria.

Discuss your project