Clinical AI Safety · Private Pilot

The Hallucination Guard
for Clinical AI.

Test clinical AI before deployment. Verify it after.

Orinyx is an independent verification platform for clinical AI. First pilot: verifying drug-to-drug and supplement-to-drug interaction recommendations against FDA labeling and clinical pharmacology.

Book a walkthrough

Orinyx is in private pilot. We're onboarding a limited founding cohort of hospitals, life-science teams, and clinical AI vendors.

  • Authoritative sources
  • Pre-deployment sandbox
  • Governance-ready reports
  • HIPAA-safe

What happens now

FIRST PILOT: MEDICATION INTERACTION SAFETY

An AI recommended aspirin 325mg daily to prevent heart attacks. It called the recommendation guideline-consistent. That guideline was reversed in 2022. The updated guidance says do not start this because bleeding risk outweighs the benefit. Nobody caught it. No record of what the AI said. No check before it reached the doctor.

31%

of US physicians use AI for clinical decisions

1–20%

medication error rate from clinical AI

$0.9M

average lawsuit from a bad AI recommendation

0 in 20

hospital AI projects that exit testing

See it in action

FIRST PILOT: MEDICATION INTERACTION SAFETY

When a recommendation isn't supported, it gets flagged. With sources.

ORINYX · CONSOLE
v0.4.2
Example verification · simulated decision dataSIMULATION
MAJOR · FLAGGED · REVIEW REQUIREDdecision #5BE2-A1
Clinician query
Patient on apixaban, can I also prescribe clarithromycin for a lung infection?
AI recommendation
"Yes, give both. Monitor INR."
Issues detected · 2
  • Fabricated monitoring: INR is only relevant for warfarin, not apixaban. The recommended monitoring parameter does not apply to this drug class.
  • Missed FDA label: Apixaban + clarithromycin requires 50% dose reduction.
    SOURCE: FDA LABEL, APIXABAN (ELIQUIS) §7.1
Sample output · this is what a flagged decision looks like in Orinyx

Want to see this run against your AI's outputs? Book a walkthrough →

One platform. Two phases.

Test clinical AI before deployment. Verify it after.

The same verification that grades an AI product before deployment keeps running once it is live. The test is the monitor.

Phase 01

Deployment Readiness

Hospitals test clinical AI products in the Orinyx sandbox before deployment. Every assertion the AI makes is verified against authoritative sources, and AI-generated clinical summaries are scored for accuracy against the source encounter. The output is a citable verification report a governance committee can act on. Phase 1 runs on synthetic data. No PHI, no EHR connection, no security review required to begin.

Phase 02

Runtime Verification

Once a product goes live, the same verification runs continuously at runtime with a full audit trail, firing upstream of the clinician. Every recommendation is checked against authoritative sources before it reaches the patient record.

In the literature

The case for independent oversight isn't ours alone.

In a prospective deployment at an academic medical center, physicians used AI-generated hospital course summaries in most discharges — and structured review of 100 summaries found omissions in 25%, inaccuracies in 20%, and one summary judged likely to cause moderate harm. The tool was useful. The review layer is what made it safe.

Ma SP, Lew T, Huynh TR, et al. "Physician-Reported Safety Outcomes of AI-Generated Hospital Course Summaries." JAMA Netw Open. 2026;9(5):e2616556. doi:10.1001/jamanetworkopen.2026.16556

Security by design

The patient record never leaves your control.

Runs in your environment

Deploy on-prem or in your private cloud. PHI stays behind your firewall.

Read-only on the record

Orinyx checks the recommendation. It never writes to the chart.

Full audit trail

Every check signed and timestamped, ready when compliance asks.

See what your clinical AI is actually saying.

Orinyx runs an independent benchmark on your AI's clinical outputs the hallucinations, misattributions, and silent drift your vendor's own monitoring won't surface. Because nothing can credibly audit itself.

We're opening a founding design-partner cohort. Founding members receive a complimentary diagnostic benchmark — paid pilots follow for teams that continue.

Frequently asked questions.