A voice agent that takes prescription orders and proves it did not mishear.
The recognizer proposes. A validator or a regulator's list decides. A refusal is a result.
A recognition error in prescription intake does not look like an error. AssemblyAI's own example: "A patient states an allergy to Lisinopril. The transcript reads Bisoprolol." Both drugs exist, both are plausible, nothing signals a fault. The vendor's own published figures set the scale: Universal-3.5 Pro Realtime has an entity error rate of 15.31%, and at 84.69% per turn a five-turn conversation "comes through clean 43.6% of the time". Reading every entity back lifts that to 79.1% in a separate vendor projection, which assumes the caller catches errors 70% of the time. A caller who hears "morphine, correct?" and says "yes" out of habit is how that assumption fails.
High recognizer confidence does not protect against homophony. The model can be certain it heard morphine while the caller said hydromorphone. Confidence proves nothing there; a regulator-published look-alike, sound-alike list does.
So a drug name on the ISMP List of Confused Drug Names is asked again even at confidence 1.00, and the question is contrastive: the agent names the drug it heard and every drug the list pairs with it, and only a spoken name answers. A reflex "yes" writes nothing.
The name is the procedure. Read-back is required of flight crews by ICAO Annex 11 §3.7.3.1 and has been a Joint Commission requirement for verbal orders since 2003, so Readback automates a step regulation already demands and busy practice skips. Why pharmacy, why this procedure and who pays for it is the product case.
|
Confidence is not evidence The pair check runs before the confidence threshold. A drug on the published list is asked about at any certainty, including 1.00, and only the caller saying a name confirms it. Naming the partner corrects the value. |
One constructor writes a field
|
Arithmetic where it exists, voice where it does not An NPI (Luhn over |
The browser holds both AssemblyAI sockets on short-lived tokens. The recognizer gives every word a millisecond span and a confidence; the agent calls our server-side tools; the gate decides what may be written.
flowchart LR
S["Caller speaks"] --> R["Recognizer<br/>words · ms spans · confidence"]
R --> P["Agent proposes a value<br/>quoting the words it came from"]
P --> G{"Gate"}
G -- "name on the ISMP list" --> C["Contrastive question<br/>only a spoken name answers"]
G -- "validator fails" --> V["Asked again,<br/>then spelled out"]
G -- "below the field's threshold,<br/>or read back by policy" --> B["Read back<br/>an explicit yes confirms"]
G -- "checksum passes,<br/>above the threshold" --> W["ConfirmedValue"]
C --> W
B --> W
V --> P
W --> O["commit_order, held:<br/>refused until every<br/>required field is proved"]
Every written field keeps its provenance: the words that produced it, their timings, the
lowest confidence among them and the validator's verdict. The order in which the gate
reads these, and the contract of every tool the agent calls, are laid out
branch by branch in the specification; how the sockets, routes,
storage and per-session agents fit together is the architecture. After the call
the server fetches the vendor's own transcript and marks each field witnessed or not in
the receipt, a second channel the browser cannot forge, described in the
security model.
The field card from the replay, under what each arm ordered: the words and timecodes behind the value, the read-back, the caller's answer and the published pair that required a spoken name.
No clone, no key and no microphone for the first four steps.
| Step | Open | What you should see |
|---|---|---|
| 1 | The replay | Hydromorphone said, morphine heard at certainty 1.00. With the pair rule on, the agent asks which of the two and the caller's "Hydromorphone." is written; with only the pair rule off, morphine is read back, a "yes" confirms it and morphine is ordered |
| 2 | Compare | Six moments side by side: said, heard, certainty, the gate's verdict, and what the same gate writes with only its threshold and validators left on |
| 3 | How it works | An attack console: seven attempts to push a value past the gate, each refused with the error the code raised, one of them before the value reaches the gate. Nothing is written |
| 4 | Metrics | Every figure with its command and set size, the pair rule's cost beside its catches, and a dash where nothing was measured |
| 5 | Call it live | Say "Hydromorphone, two milligrams". The agent names the drug and every drug the list pairs with it, and waits for a name |
What each step proves, why each attack fails and where every judging criterion is answered are set out in the guide for judges.
Step 1, the replay. The same synthesised session through the shipped gate twice, differing by one flag.
The receipt of a production call placed by the synthesised-caller smoke harness, rechecked in the browser: the sha256, the DEA and NPI check digits and the catalogue.
Every figure below is printed by a command in this repository, offline and free. The measurement report carries each one with its command and set size, and the evidence page grades every claim the project makes, from measured to assumed.
| Figure | What it shows | Command |
|---|---|---|
| Without the pair rule a reflex yes writes 20 of 20 pair mishearings; with it, 0 | The pair rule is the only difference between the two arms | npx tsx scripts/measure/ab-gate.ts |
| 59 of 59 correct drug names are asked about: 25 by the standing read-back, 13 by the threshold, 21 by a contrastive question | The cost side: which question a correct value gets | npx tsx scripts/measure/coverage-matrix.ts |
| The 2023 ISMP list carries 514 pairs over 754 names; 502 of 3730 catalogue drugs carry a listed name | How much of the catalogue the pair rule reaches | npx tsx scripts/measure/ismp-coverage.ts |
| NPI catches 100% of substitutions, DEA catches 95.2%, over 32 080 mutations | The two checksums are not equally strong, and the policy says so | make audit-checksums |
| 12 of 12 mutations killed | Every gate branch fails its own named test when broken | make gate-mutation |
The 20 curated pairs, each checked by hand against its row of the list, drive the demo and the evaluation; the product rule reads the whole list.
npm ci
make data
make dev
make verifymake dev creates .env.local from .env.example when it is missing and serves
http://localhost:3000. The tests, the replay and make honest (every offline figure with its
command) need no key. A live call needs ASSEMBLYAI_API_KEY, AGENT_TOOL_SECRET and a public
HTTPS origin in .env.local, because AssemblyAI calls the agent's tools only on a public host:
a laptop needs a tunnel, and otherwise the tools are exercised on a deployment. Every variable
is listed in how to run it locally. make verify
runs every check; what each one refuses, and which few commands spend credit, is in
how it is verified.
| Capability | Why the product depends on it |
|---|---|
Word-level timings and per-word confidence (universal-3-5-pro) |
The gate takes the lowest confidence over the words behind a value; a turn-level score would hide the one failed word that is the drug name |
| Voice Agent API with server-side HTTP tools | AssemblyAI itself calls propose_field, read_back and commit_order on our server, so each tool call is a short HTTPS request to a serverless route and none runs in the browser |
commit_order in hold execution mode |
The agent waits for the server's verdict instead of talking over a write |
keyterms_prompt |
Biases the recognizer toward clinic, prescriber and unit words, and is forbidden, by a check, from ever containing a drug name the pair rule tests |
| Entity-aware turn detection | The vendor documents that the agent waits for a whole entity before ending a turn; we never send the two settings that switch it off, because an NPI dictated in digit groups is the case it protects |
- The gate proves provenance, not truth. Provenance is computed in the browser, so a hostile client can report words nobody said. The vendor's own transcript, fetched by the server after the call, witnesses each field, but it cannot block a commit during the call. The threat model states what is and is not protected.
- The pair rule is exactly as good as the 2023 ISMP list. AssemblyAI's own example, lisinopril heard as bisoprolol, is on no published list, so only the plain read-back, which a reflex "yes" passes, stands between it and the order.
- The evaluation audio is synthesised. Recognition is real AssemblyAI traffic over desktop text-to-speech voices; no human-voice set has been measured. The recognizer has not been caught turning a listed name into its published partner in our corpus (0 of 186 degraded utterances), so the rule rests on the published list, not on our data.
- Live calls on production have run with a synthesised caller, not yet a human voice.
These are the short versions. Each one, with what is measured, what enforces it and what remains open, is set out in full in the limitations.
A technology demonstration, not a medical device. Synthetic data only: no real patients, no real prescriptions. In an emergency call 911, or 988 for a mental health crisis.
Built for the AssemblyAI Voice Agent Hackathon.



