Workflow / MLR Review
Flag, Judge, Prove is the three-stage workflow for AI-assisted MLR review: the AI flags candidate issues, a human reviewer judges each one, and the system proves the call with a signed audit record. The hard rule that holds it together: no LLM sits inside the flag decision plane. The doctrine is "It flags. You decide." — the model assists, and a named person makes every determination.
MLR — Medical, Legal, and Regulatory review — is the gate pharmaceutical promotional material passes through before it reaches the public. It is slow, high-stakes work: a missed unsubstantiated claim or a misstated fair balance is a regulatory liability, not a typo. The temptation is to hand the judgment to a model. The Flag → Judge → Prove workflow is a deliberate refusal to do that. It uses AI to make the review faster to start without letting AI make the review's decisions.
The assistant reads the promotional material and surfaces candidate issues — a claim that may need substantiation, a place where fair balance looks thin, a citation that may not support the statement it is attached to. These are candidates for review, not verdicts. The output of this stage is a shortlist that tells a busy reviewer where to look first, nothing more.
A qualified reviewer works each flag and makes the call: accept it, dismiss it, or edit the material. This is where the decision lives, and it lives with a person. The reviewer can act on a flag, overrule it, or find something the system missed. The AI does not approve, reject, or clear anything. It never had that authority.
Each flag and the human judgment on it is written to a tamper-evident audit chain and signed with post-quantum cryptography (ML-DSA). The record captures what was flagged, what the reviewer decided, and when — in a linked sequence whose integrity anyone with the public key can verify. The decision is not just made; it is provable afterward.
This is the line the whole workflow is built to protect. An LLM can help surface candidates, but it is never wired into the plane where a flag becomes a decision. The reason is accountability. A language model produces plausible text; it does not carry professional responsibility, cannot be deposed, and cannot sign its name to a determination. In a regulated review, someone has to own the call. That someone is the reviewer.
Keeping the model out of the decision plane also keeps the audit record honest. When the person who decided is always human and always named, "who made this call" has a real answer every time. That is a very different posture from a system that lets a model auto-approve the easy cases and asks a human to review only the hard ones — because in that design, no one can say cleanly where the machine stopped and the person started.
The Prove stage is what makes the workflow defensible rather than merely careful. Each flag and each human judgment becomes a record, each record is chained to the one before it, and the chain is covered by a digital signature. Change an earlier decision, insert a call that never happened, or drop one that did, and the signature no longer validates and the chain no longer links. Tampering does not stay hidden; it surfaces the moment anyone verifies.
The signatures use ML-DSA, the NIST-standardized post-quantum digital signature algorithm, so a record you must retain for years does not rest on a scheme a future quantum computer could forge. And verification does not depend on the vendor: anyone holding the public key and the chain can confirm, on their own hardware, that the record is intact. For the mechanics, see signed, auditable AI.
The signed audit chain that powers Prove is the fourth of four controls in a Hyfstele deployment, each verifiable on your own hardware, inside your own perimeter — open-weight models, air-gapped:
For a life-sciences buyer this is not a philosophical preference; it is what makes the process inspectable. 21 CFR Part 11 expects secure, time-stamped audit trails that record who did what and when and do not obscure prior information, with records kept attributable and protected against alteration. GxP expects decisions traceable to a responsible person. Data-residency and model-provenance requirements expect you to know where the model runs and prove it is the model you vetted.
Flag → Judge → Prove maps onto those expectations directly. Judge keeps every determination attributable to a named human. Prove makes each determination time-stamped, linked, and tamper-evident. And the no-LLM-in-the-decision-plane rule means the audit trail never has to explain away a machine that decided on its own. When a reviewer, an auditor, or an inspector later asks "why was this claim flagged, and who cleared it," the answer is a single signed record — not a best-effort reconstruction after the fact.
Hyfstele's MLR use case is a live AI assist for pharmaceutical promotional review, running the Flag → Judge → Prove flow with no LLM in the flag decision plane. You can try the working demonstration at mlr.hyfstele.com. For the wider context, see AI for pharma MLR review and MLR review software.
It is a three-stage workflow for AI-assisted MLR (Medical, Legal, Regulatory) promotional review. In Flag, the system surfaces candidate issues for a reviewer to look at. In Judge, a human reviewer decides to accept, reject, or edit each flag. In Prove, every flag and every human judgment is written to a tamper-evident audit chain signed with post-quantum cryptography (ML-DSA), so the record stands as evidence later.
No. The hard rule is that there is no LLM inside the flag decision plane. The doctrine is "It flags. You decide." The AI surfaces candidates worth a human look, but the human reviewer holds every decision. This keeps accountability with a qualified person and keeps the model assisting rather than deciding.
It means the AI's job ends at surfacing candidates, and the reviewer's job is the decision. The assistant never approves, rejects, or clears promotional material on its own. A named person makes every determination and owns it, which is what a regulated review process requires.
Every flag and every human judgment is anchored to a tamper-evident audit chain signed with ML-DSA, a post-quantum signature standard. Because the signature covers each linked record, any later alteration, insertion, or deletion is detectable and independently verifiable. When someone later asks why an item was flagged and who cleared it, the answer is a signed record, not a reconstruction.
There is a live demonstration of the MLR assist at mlr.hyfstele.com. It runs the Flag → Judge → Prove flow with no LLM in the flag decision plane, inside a Hyfstele deployment where open-weight models run air-gapped within your own perimeter.
Hyfstele runs open-weight models inside your perimeter, air-gapped, with four controls you can verify on your hardware. Try the MLR assist live, or talk through a deployment.
Open the live demo Email Blake →