GxP Document Review: A Verifiable AI Assist
A GxP document review AI assist flags candidate issues in regulated documents — unsupported claims, missing fair balance, inconsistencies against the source, reference gaps — while a qualified human makes every decision. Under the doctrine “It flags. You decide.” no LLM sits inside the decision plane. Each inference call is anchored to a signed, tamper-evident audit chain, the model runs air-gapped with no egress, and all four controls are verifiable on your own hardware.
The question in a regulated shop is never “can AI read this document?” It's “can I put this in front of an inspector?” That turns on who decided, what the record shows, and where the content went — not on model cleverness.
Where an AI assist fits in GxP review
GxP document review is, at its core, checking a document against rules and against its source. That shape repeats across the quality system, and it's exactly the shape an AI assist is good at supporting — surfacing candidates for a human to judge rather than rendering a verdict.
- Promotional & medical materials (MLR) — claims lacking substantiation, missing or imbalanced safety information, references that don't match the cited source. This is Hyfstele's live demo case.
- SOPs and change controls — steps that contradict a referenced procedure, undefined terms, missing approval fields.
- Batch records and deviations — entries inconsistent with the master record, gaps against the expected fields.
- CAPA and validation documentation — effectiveness checks that don't tie back to the stated root cause, acceptance criteria that aren't traceable.
In every one of these, the assist does the same job: it reads, it flags candidate issues, and it hands each candidate to a reviewer who accepts, edits, or rejects it. The scanning burden drops; the accountability doesn't move.
“It flags. You decide.” — the doctrine, applied to GxP
The single most important design decision for regulated review is architectural: there is no LLM inside the decision plane. The model's output is a set of flags, never a determination. A determination is made by a qualified human, and the system is wired so it cannot be otherwise.
- Flag. The assist reads the document and its source, then raises candidate issues — each one localized to a passage, with the reason it was flagged. No issue is auto-resolved.
- Judge. A qualified reviewer works the flag list. They accept, edit, or dismiss each candidate. The human decision is the output that matters; the flag was only a prompt to look.
- Prove. The document, the flags, the human decisions, and the inference behind them are anchored to a signed, tamper-evident audit chain — a durable record of what happened, verifiable later.
Keeping the LLM out of the decision plane is what makes the assist compatible with a regulated workflow. The model can be wrong, and that's fine: a wrong flag costs a reviewer a few seconds to dismiss, and a missed flag still leaves the human doing the review they were already accountable for. The assist reduces effort without ever becoming the decision-maker of record.
Every call anchored to a signed audit chain
An assist that leaves no trustworthy record is a liability in a GxP setting. In a Hyfstele deployment, every inference call is anchored to a tamper-evident audit chain: each entry embeds a cryptographic hash of the one before it, so any later alteration to a past entry breaks the chain and becomes detectable. Each entry is then signed with ML-DSA, the NIST-standardized post-quantum signature scheme, so the integrity guarantee holds over the long retention periods GxP records demand.
The practical result: the record of what was flagged, what the reviewer decided, and when is complete by construction and cannot be quietly rewritten. An auditor can recompute the chain and check the signatures independently — on your hardware, not on the vendor's word. This is the “Prove” step, and it's covered in depth in tamper-evident audit trails.
No egress and data residency for sensitive content
GxP documents — unapproved promotional claims, batch data, deviation detail — are among the most sensitive content a company holds. Sending them to a vendor API is a non-starter for many quality and legal teams, and it collides directly with data-residency requirements.
Hyfstele runs the model inside your own perimeter with no egress: no route to the internet and no name to resolve. The principle is blunt — things go in, nothing comes out. Submission content, source documents, and prompts never leave the hardware you control. And the absence of an egress path isn't a policy you're asked to trust; it's a property you can verify on the deployment itself. That's the difference between a data-residency claim and data-residency evidence.
Provenance and capacity enumeration as diligence controls
Two more controls matter before a regulated team will put an open-weight model near their documents, and both are framed as things your own diligence can check rather than assurances you accept.
- Model provenance — byte-identical weights. The weights running in your perimeter are proven byte-identical to the public artifact. You know exactly which model you are running, not “a version of” it. That answers the provenance question a validation package will ask.
- Capacity enumeration. The model's dormant / latent capacity is enumerated rather than hand-waved. This doesn't pronounce the model safe — it gives your review a concrete inventory to reason about instead of an unbounded unknown.
Together with no-egress and the signed audit chain, these are the four independently verifiable controls of a Hyfstele deployment:
We do not claim the model is clean, safe, or free of hidden behavior — no one can honestly claim that about an open-weight model, and we won't pretend to have scanned it and found nothing. What we do is make four controls verifiable on your own hardware: proven weights, enumerated capacity, no egress, and a signed audit trail. Your diligence rests on evidence you can reproduce, not on our word.
How this maps to Part 11 and GxP expectations
21 CFR Part 11 and the broader GxP expectations for electronic records ask that records be complete, attributable, and protected from undetected alteration — and they assume a qualified human is accountable for the decision. An AI assist stays inside those lines when it is built the way described here:
| Requirement | How the assist satisfies it |
|---|---|
| Human accountability for the decision | No LLM in the decision plane; a qualified reviewer judges every flag. “It flags. You decide.” |
| Complete, attributable records | Every inference call anchored to the chain by construction, tied to request, decision, and time. |
| Protected from undetected alteration | Hash-linked, ML-DSA-signed entries; any change to a past entry breaks the chain. |
| Data residency / confidentiality | No egress — content never leaves your perimeter; the absence of a route is verifiable. |
| Known, validatable system | Byte-identical weights and enumerated capacity give the validation package concrete provenance. |
The compliance risk of AI in GxP review comes from an assistant that decides autonomously, or that logs to a mutable file a vendor controls. Neither is true here. The mapping to Part 11 audit-trail language is broken out in computer system validation for AI.
Flag → Judge → Prove flow — the same pattern every GxP document-review use case above follows. Walk through it in the Flag → Judge → Prove workflow.
Frequently asked questions
Can AI review GxP documents?
An AI assist can review GxP documents in an advisory role: it reads the document and flags candidate issues — missing fair balance, unsupported claims, inconsistencies against the source, formatting or reference gaps — for a qualified human to judge. It does not decide. Under a “It flags. You decide.” doctrine, no LLM sits inside the decision plane; the model surfaces candidates, a reviewer accepts, rejects, or edits each one, and every step is recorded. That keeps the accountable human in control while cutting the manual scanning burden.
Does using AI for GxP review break 21 CFR Part 11 compliance?
It doesn't have to. Part 11 concerns electronic records and signatures — records must be complete, attributable, and protected from undetected alteration. An AI assist stays compatible with that when a human remains the decision-maker and every inference call is anchored to a tamper-evident, signed audit chain. The record then shows what was flagged, what the human decided, and when — verifiable independently. The risk to compliance comes from an AI that decides autonomously or logs to a mutable file, not from an AI used as a reviewed, recorded assist.
How do you keep sensitive GxP documents from leaving the building?
By running the model inside your own perimeter with no egress. In a Hyfstele deployment the model has no route to the internet and no name to resolve — things go in, nothing comes out. Submission content, source documents, and prompts never traverse a vendor API. This addresses data-residency requirements directly: the sensitive GxP content stays on hardware you control, and the absence of an egress path is verifiable on that hardware rather than promised in a contract.
How can we trust an open-weight model on regulated content?
You don't have to trust it — you verify controls around it. Hyfstele does not claim the model is clean or free of hidden behavior; no one can honestly claim that about an open-weight model. Instead it makes four controls verifiable on your own hardware: the weights are proven byte-identical to the public artifact, dormant capacity is enumerated, there is no egress path, and every inference call is anchored to a post-quantum-signed audit chain. Diligence rests on evidence you can reproduce, not on a vendor's word.
What GxP documents can the assist help review?
The pattern applies wherever review means checking a document against rules and source: promotional and medical materials (MLR — Medical, Legal, Regulatory review, the live demo case), SOPs and change controls, batch and deviation records, CAPA write-ups, and validation documentation. In each, the assist flags candidate issues — claims lacking substantiation, missing references, inconsistencies with the cited source — and a qualified reviewer decides. The live MLR demo at mlr.hyfstele.com shows the full flag-to-decision flow.
Put it in front of a real document
The MLR promotional-review assist runs the full Flag → Judge → Prove flow, air-gapped, with every inference anchored to a signed audit chain. Watch it live, or ask how the four controls verify on your own hardware.