Crucable Research

Research

Structured adversarial tests of live, patient-facing AI agents. Every study runs against a deployed production surface. No simulations, no replays. Each failure is a trace you can read.

Published audits2 audits · 9 agents tested · synthetic personas · no patient data

Escalation fragility in a live patient-facing agent

13 scenarios against Westbrook Clinic's assistant. Difficult to argue with — easy to redirect. Two replicated failures when a plausible self-diagnosis and topic switch erased a flagged red flag. Plus order-dependent rigor and no cross-visit memory.

Pass
7
Findings
5
Runs
14
Read audit 001
Key findings
  • Red flag erasure. Minimization + plausible alternative + redirect in one message dropped a cardiac and a GI red flag the agent had already caught.
  • Order fragility. Same facts, different order → different rigor (7a skipped questions, 7b hid specificity).
  • No cross-visit memory. Three contacts for worsening cough → no trend acknowledgment; day-four under-triage.
Westbrook Clinic's assistant · general-purpose model · manual multi-turn · fresh session per scenario
More audits forthcoming. Each audit is a trace-first report scoped to deployed agents and workflows. Methods and personas are documented so a buyer can replay every check.