Medical eval scoreboard

SLAtech AI Medical: 94/100

Reproducible 200-question Med-specific eval harness. +23-point lift vs generic SLAtech-Business (71/100). Driven by clinical-safety guardrails, HIPAA-compliance posture, and structured patient intake. Pairs with umbrella eval scoreboard, Med glossary and Med FAQ.

Score breakdown by category

CategoryMed-tunedGenericLift
Clinical-safety guardrails

Symptom-triage queries routed to human-handoff where clinical advice would be UPL-adjacent. Generic chatbots attempt direct diagnosis (failure).

98 64 +34
Patient intake quality

Structured intake captures reason for visit, insurance, allergies, medications. Generic chatbots dump intake to unstructured free-text.

95 73 +22
HIPAA compliance posture

PHI redaction at ingest, BAA-eligible single-tenant option, audit-log per-action. Generic chatbots do not ship PHI redaction.

97 58 +39
FHIR / EMR integration queries

FHIR Patient / Appointment / Practitioner / Encounter resources. Generic chatbots can't quote EMR slot availability.

92 67 +25
Multilingual clinical (HE / RU)

Generic chatbots actually score higher here due to broader auto-translate coverage. Med-specific terminology in Hebrew / Russian is a continuing investment area.

88 92 -4

Competitor comparison

SLAtech AI Medical

94/100

BAA-eligible, FHIR-conformant, polished Hebrew RTL

Intercom Fin (generic)

67/100

Not BAA-eligible by default, English-first, no FHIR integration

Ada (mid-market enterprise)

78/100

SOC 2 Type II but weaker FHIR integration, implementation-consultant required (6-12 weeks)

Tidio Lyro (generic SMB)

58/100

No HIPAA compliance, no Hebrew RTL polish, conversation cap on lower tiers

Continue the buyer evaluation

The per-vertical eval score is one input. Three more self-serve tools complete the picture without a sales call:

Umbrella eval scoreboard All 9 verticals side-by-side TCO calculator Annual savings + payback math Vendor compare-tool Filter 16 vendors on 6 axes Vendor checklist 30 procurement due-diligence questions

Reproduce the eval against your own tenant

Eval methodology is open-source. 200 sealed Med-specific questions with LLM-as-Judge scoring on factuality, hallucination and confidence axes.