Workflow / MLR Review

The Flag → Judge → Prove Workflow for MLR Review

Flag, Judge, Prove is the three-stage workflow for AI-assisted MLR review: the AI flags candidate issues, a human reviewer judges each one, and the system proves the call with a signed audit record. The hard rule that holds it together: no LLM sits inside the flag decision plane. The doctrine is "It flags. You decide." — the model assists, and a named person makes every determination.

MLR — Medical, Legal, and Regulatory review — is the gate pharmaceutical promotional material passes through before it reaches the public. It is slow, high-stakes work: a missed unsubstantiated claim or a misstated fair balance is a regulatory liability, not a typo. The temptation is to hand the judgment to a model. The Flag → Judge → Prove workflow is a deliberate refusal to do that. It uses AI to make the review faster to start without letting AI make the review's decisions.

Flag Judge Prove

The three stages

  1. Flag Surface the candidates

    The assistant reads the promotional material and surfaces candidate issues — a claim that may need substantiation, a place where fair balance looks thin, a citation that may not support the statement it is attached to. These are candidates for review, not verdicts. The output of this stage is a shortlist that tells a busy reviewer where to look first, nothing more.

  2. Judge The human decides

    A qualified reviewer works each flag and makes the call: accept it, dismiss it, or edit the material. This is where the decision lives, and it lives with a person. The reviewer can act on a flag, overrule it, or find something the system missed. The AI does not approve, reject, or clear anything. It never had that authority.

  3. Prove Signed evidence of the call

    Each flag and the human judgment on it is written to a tamper-evident audit chain and signed with post-quantum cryptography (ML-DSA). The record captures what was flagged, what the reviewer decided, and when — in a linked sequence whose integrity anyone with the public key can verify. The decision is not just made; it is provable afterward.

The hard rule: no LLM in the flag decision plane

This is the line the whole workflow is built to protect. An LLM can help surface candidates, but it is never wired into the plane where a flag becomes a decision. The reason is accountability. A language model produces plausible text; it does not carry professional responsibility, cannot be deposed, and cannot sign its name to a determination. In a regulated review, someone has to own the call. That someone is the reviewer.

The rule, stated plainly "It flags. You decide." The AI's role ends at surfacing candidates worth a human look. Every accept, reject, or edit is a human act. There is no path in the system where a model clears promotional material on its own — not as a default, not as a fast-track, not as a fallback.

Keeping the model out of the decision plane also keeps the audit record honest. When the person who decided is always human and always named, "who made this call" has a real answer every time. That is a very different posture from a system that lets a model auto-approve the easy cases and asks a human to review only the hard ones — because in that design, no one can say cleanly where the machine stopped and the person started.

How every call is anchored to a signed audit chain

The Prove stage is what makes the workflow defensible rather than merely careful. Each flag and each human judgment becomes a record, each record is chained to the one before it, and the chain is covered by a digital signature. Change an earlier decision, insert a call that never happened, or drop one that did, and the signature no longer validates and the chain no longer links. Tampering does not stay hidden; it surfaces the moment anyone verifies.

Flag raised Human judgment Linked record ML-DSA signature

The signatures use ML-DSA, the NIST-standardized post-quantum digital signature algorithm, so a record you must retain for years does not rest on a scheme a future quantum computer could forge. And verification does not depend on the vendor: anyone holding the public key and the chain can confirm, on their own hardware, that the record is intact. For the mechanics, see signed, auditable AI.

Honest posture We do not claim the model is clean, safe, or free of hidden behavior — no scan can prove that away. What we claim is narrower and checkable: the audit chain proves the record of what happened is intact, and it is one of four controls you can verify on your own hardware. Your confidence comes from evidence you can reproduce, not from our word.

One control inside four

The signed audit chain that powers Prove is the fourth of four controls in a Hyfstele deployment, each verifiable on your own hardware, inside your own perimeter — open-weight models, air-gapped:

  1. Weights proven byte-identical to the public artifact, so the model you run is the model you audited.
  2. Dormant / latent capacity enumerated, so unused capability is accounted for rather than assumed away.
  3. No egress — no route to the internet and no name to resolve. Things go in; nothing comes out.
  4. Every inference call anchored to a tamper-evident audit chain signed with post-quantum cryptography (ML-DSA). This is what the Prove stage writes to.

Why this separation matters for regulated accountability

For a life-sciences buyer this is not a philosophical preference; it is what makes the process inspectable. 21 CFR Part 11 expects secure, time-stamped audit trails that record who did what and when and do not obscure prior information, with records kept attributable and protected against alteration. GxP expects decisions traceable to a responsible person. Data-residency and model-provenance requirements expect you to know where the model runs and prove it is the model you vetted.

Flag → Judge → Prove maps onto those expectations directly. Judge keeps every determination attributable to a named human. Prove makes each determination time-stamped, linked, and tamper-evident. And the no-LLM-in-the-decision-plane rule means the audit trail never has to explain away a machine that decided on its own. When a reviewer, an auditor, or an inspector later asks "why was this claim flagged, and who cleared it," the answer is a single signed record — not a best-effort reconstruction after the fact.

See it working

Hyfstele's MLR use case is a live AI assist for pharmaceutical promotional review, running the Flag → Judge → Prove flow with no LLM in the flag decision plane. You can try the working demonstration at mlr.hyfstele.com. For the wider context, see AI for pharma MLR review and MLR review software.

Frequently asked questions

What is the Flag, Judge, Prove workflow for MLR review?

It is a three-stage workflow for AI-assisted MLR (Medical, Legal, Regulatory) promotional review. In Flag, the system surfaces candidate issues for a reviewer to look at. In Judge, a human reviewer decides to accept, reject, or edit each flag. In Prove, every flag and every human judgment is written to a tamper-evident audit chain signed with post-quantum cryptography (ML-DSA), so the record stands as evidence later.

Does an LLM make the MLR decision in this workflow?

No. The hard rule is that there is no LLM inside the flag decision plane. The doctrine is "It flags. You decide." The AI surfaces candidates worth a human look, but the human reviewer holds every decision. This keeps accountability with a qualified person and keeps the model assisting rather than deciding.

What does "It flags. You decide." mean?

It means the AI's job ends at surfacing candidates, and the reviewer's job is the decision. The assistant never approves, rejects, or clears promotional material on its own. A named person makes every determination and owns it, which is what a regulated review process requires.

How is each decision made defensible?

Every flag and every human judgment is anchored to a tamper-evident audit chain signed with ML-DSA, a post-quantum signature standard. Because the signature covers each linked record, any later alteration, insertion, or deletion is detectable and independently verifiable. When someone later asks why an item was flagged and who cleared it, the answer is a signed record, not a reconstruction.

Where can I see the workflow working?

There is a live demonstration of the MLR assist at mlr.hyfstele.com. It runs the Flag → Judge → Prove flow with no LLM in the flag decision plane, inside a Hyfstele deployment where open-weight models run air-gapped within your own perimeter.

See Flag → Judge → Prove in action

Hyfstele runs open-weight models inside your perimeter, air-gapped, with four controls you can verify on your hardware. Try the MLR assist live, or talk through a deployment.

Open the live demo Email Blake →