Skip to main content
All articles

Blog

Where AI Should Stop in a Regulated SAP Workflow: Confidence Thresholds, Specification Limits, and Expert Judgment in Pharma and Aerospace QM

This article explores automating quality management in regulated industries (pharma, aerospace), highlighting the need to balance AI document processing with human oversight for compliance.

Diagram illustrating SAP process automation driven by an AI-powered platform.

A pallet of micronized API arrives at the dock at 6:40 on a Tuesday morning. The goods receipt posts, the stock lands in quality inspection, and an inspection lot opens against batch 24-0817. Stapled to the shipping documents is a three-page certificate of analysis from the supplier. Assay 99.2 percent against a specification of 98.0 to 102.0. Water by Karl Fischer 0.48 percent against a limit of not more than 0.50. Residual acetone 412 ppm against a limit of 500. Particle size D90 at 8.7 microns inside a window of 5 to 10.

An extraction model reads that certificate in about four seconds and reports high confidence on every field. Every number is inside its limit. The lot could, in principle, move to unrestricted stock before the quality engineer has finished her first coffee.

That is exactly the moment where the interesting question starts. Not "can the system read the certificate," because it can. The question is which of the decisions sitting downstream of that certificate the system is allowed to make, and which ones belong to a named human being who will sign for them. Get that line in the wrong place and you have not automated a quality process. You have automated your way into a finding.

Three different questions that get treated as one

Most conversations about AI in quality management collapse into a single metric, usually phrased as accuracy. A vendor says the model is 99.4 percent accurate, a quality manager asks whether that is good enough, and everyone argues about a number that does not answer the question being asked.

There are actually three questions stacked on top of each other, and only the first one is a machine learning problem.

The first question is "did I read it right." The certificate says 0.48 and the system recorded 0.48. This is transcription. A confidence score speaks to this and nothing else.

The second question is "is this value acceptable." The recorded 0.48 sits against a limit of 0.50. This is arithmetic plus metrology, and it depends on rounding conventions, unit handling, the measurement uncertainty of the test method, and whether the number reported is an individual result or a mean of three.

The third question is "may this material be used." That is a disposition, and in a regulated plant it is a legal act performed by a person with specific delegated authority. In pharmaceutical manufacturing, US regulation places the authority to approve or reject components and finished product squarely with the quality control unit, and in the European system a Qualified Person certifies each batch personally before it goes anywhere near a patient. In aerospace, the authority to accept something that does not meet the drawing sits with the design authority or the customer, not with the supplier who made it.

A confidence score of 99.4 percent has nothing to say about questions two and three. Treating it as though it does is the single most common design error in regulated document automation.

What the regulators have now put in writing

For years the guidance on automated decisions in GxP systems had to be inferred from computerised systems rules written before modern models existed. That gap is closing fast.

The European Medicines Agency and the European Commission put a draft Annex 22 on artificial intelligence out for consultation, with the comment period closing in October 2025 and finalisation expected during 2026. It arrived alongside companion revisions to Annex 11 on computerised systems and Chapter 4 on documentation, which tells you the regulators see AI as an extension of existing system validation rather than a separate universe.

The draft is unusually direct about scope. It covers static models whose parameters do not change during use. Models that keep learning in production are not acceptable for critical GMP applications, and neither are generative models or large language models where a GMP-critical decision depends on the output. Those same generative tools are fine for lower-stakes work like drafting a maintenance note or summarising a deviation report, as long as a qualified person reviews the output and keeps documented responsibility for it.

Three requirements in that draft matter enormously for how you design a document workflow. Intended use has to be defined and approved before testing begins, including the edge cases and known failure modes. The model has to perform at least as well as the process it replaces, measured against metrics agreed in advance. And confidence scores have to be logged, with the model returning an explicit "undecided" when confidence falls very low rather than guessing.

That last point is worth sitting with. The regulator is not asking for a higher confidence threshold. The regulator is asking for a system that knows when to stop.

On the US side, the FDA's draft guidance on AI supporting regulatory decision-making, issued in January 2025, builds a seven-step credibility framework around the same instinct. You define the question of interest, define the context of use, then assess risk along two axes. One axis is model influence, meaning how much the model's output actually drives the decision. The other is decision consequence, meaning how bad it is if the output is wrong. A model that flags items for a human to re-inspect is low risk. A model that decides on its own, where being wrong reaches a patient, is high risk and needs the strongest evidence you can produce.

Notice what the FDA framework gives you. If your credibility evidence is not strong enough for the role you wanted the model to play, one of the listed remedies is simply to reduce the model's influence. Move it from deciding to proposing. That is a design lever, not a failure.

Aerospace has no single AI annex yet, and the pressure arrives through a different door. AS9100 requires that authority to accept nonconforming product comes from the customer where the customer reserves it. AS9102 governs first article inspection and the evidence that a production process actually makes what the drawing says. AS13100, the aero engine supplier standard published through the AESQ and revised in 2025, layers on defect prevention tooling including a measurement systems analysis reference manual. None of these documents mention neural networks. All of them constrain what an automated system may conclude on its own, because they define whose signature carries weight.

Diagram showing the Four Zones of Decision Authority, mapping organizational roles to levels of decision-making power.

Why a single confidence threshold is the wrong control

The instinct once you accept that confidence matters is to set a number. Ninety-five percent auto-posts, below that goes to review. It feels rigorous. It is close to useless on its own, for four reasons that show up in real document sets within weeks.

Confidence is not accuracy. A model's confidence is its own estimate of how likely it is to be right, and that estimate can be badly calibrated. A model may be 95 percent confident across a thousand fields and correct on only 88 percent of them. Until you have measured calibration against a challenge set drawn from your actual supplier base, the threshold you picked is a guess wearing a number.

Confidence is field-level, and risk is not evenly spread. The same certificate carries a storage condition in free text, a batch number, and a water content result. Misreading the storage condition is an annoyance. Misreading a digit in the batch number breaks traceability across every downstream record. Misreading the water content can release a material that should have been rejected. One threshold across all three fields either over-reviews the harmless fields or under-reviews the dangerous ones. Thresholds belong to characteristics, not to documents.

High confidence and wrong is the failure that actually hurts. Picture a supplier who reformats their certificate template and moves the specification column next to the result column. The model, which learned that the second numeric column on this layout is the result, reads the specification value instead. It reports 98.0 with 98 percent confidence. It is not hallucinating and it is not uncertain. It read a real number, printed clearly, in the wrong place. The true assay was 97.4 and the batch should have failed. No confidence threshold catches this, because the model is confident and, in its own terms, correct.

What catches it is a different control entirely. Cross-field consistency checks, a specification column that is read and compared against the specification held in the system rather than ignored, and layout change detection that raises a flag when a supplier's document no longer matches the fingerprint on file. Those controls are cheap. They are also the ones most often left out, because they are not glamorous and no dashboard reports them.

Confidence says nothing about semantics. A quantitative characteristic expects a number. The certificate says "Conforms," or "ND," or "less than 0.05." A residual solvent arrives as 0.041 percent from one supplier and 412 ppm from another, against a limit written in ppm. Recorded literally, 0.041 against a limit of 500 looks gloriously in specification. It is the same value as 410 ppm, and the arithmetic that would have caught it is a unit conversion the extraction step never performed. The model was right about every character on the page.

Specification limits are not decision rules

Here is where quality engineers and automation vendors talk past each other most reliably. A specification limit is a boundary. A decision rule is the statement of what you do with a measured value near that boundary, and the two are not the same document.

Take the water content at 0.48 against a limit of 0.50. Whether that passes depends on questions the number alone does not answer. Is 0.48 a single determination or the mean of duplicates that read 0.44 and 0.52? Under most valuation setups, a mean comparison accepts the pair, and an individual value outside the limit disappears into the average. What is the expanded uncertainty of the Karl Fischer method as your lab runs it? If it is plus or minus 0.03 at the coverage factor you use, then 0.48 and 0.50 are not distinguishable with confidence, and a laboratory operating under a documented decision rule of the kind described in the ILAC guidance on conformity statements would apply a guard band rather than call a clean pass. What rounding convention applies? A value of 0.4951 reported to two decimals becomes 0.50, which is at the limit rather than under it, and the pharmacopoeial convention is to round to the same number of decimal places as the specification before comparing.

None of that is exotic. All of it is invisible to a model that compares two numbers.

Aerospace lives the same problem in metal. A bore on a compressor disc measures 25.012 millimetres against a drawing of 25.000 plus 0.015 minus 0.000. Inside tolerance by inspection, comfortably. Now ask what the gage repeatability and reproducibility is for that measurement. If the measurement system consumes 30 percent of the tolerance band, the instrument cannot resolve the difference between a good part and a marginal one, and AS13100's measurement systems analysis manual exists precisely to force that question before anyone trusts the reading. The number on the report is not wrong. It just carries less information than it appears to.

For a key characteristic under statistical control the bar rises again. Conformance of the individual part is not the whole answer, because the standard asks about the capability of the process producing it. A part at the top of tolerance from a process running a capability index below the required threshold is a signal about the next hundred parts, not just this one. An automated conformance check that says "pass" on the individual reading is technically correct and strategically blind.

The practical consequence for system design is a band, not a line. Every quantitative characteristic needs three zones rather than two. There is a zone where the value is inside the limit by more than the combined uncertainty and rounding margin, which can be auto-valuated. There is a zone where the value is outside the limit by more than that margin, which is a clear rejection and an automatic trigger for investigation. And there is a zone in between, the guard band, where a machine may record and present the value but may not conclude anything about it. That middle zone is where your most experienced people earn their salary, and the system's job is to route work to them rather than to resolve it.

Walking the stop points through a live inspection lot

Abstract principles get argued about. A specific lot moving through a specific workflow gets decided. So walk the API batch and the titanium forging through the same sequence and mark where the machine hands over.

Lot creation and specification retrieval are safe territory. The goods receipt posts, the inspection type on the material triggers a lot, the inspection plan or material specification attaches, and sample size calculates. No judgment involved. This should be fully automatic and mostly already is.

Reading the incoming document is safe territory too, and this is where the real time is won. A three-page certificate with eighteen characteristics, a heat treatment certificate, a material test report, a first article inspection report with 214 characteristics across twelve pages. Extracting those values, normalising units, matching them to the right master inspection characteristic, and checking the document is actually for the batch and material on the purchase order is work a machine does faster and more consistently than a person at 6:40 in the morning. The failure mode of a human here is not misjudgment. It is transcription drift on page nine.

Recording results into the inspection lot is safe, with one condition. The system writes what it read, it writes the source it read it from, and it writes its confidence alongside. Where confidence falls below the threshold for that characteristic, it writes nothing and marks the characteristic as requiring entry. An empty field a human must fill is a good outcome. A guessed field is not.

Valuation of the characteristic is where the first real boundary sits. Comfortably inside the limits, automatic valuation is reasonable and defensible, because the logic is deterministic arithmetic against a stored specification and it is testable to exhaustion. Inside the guard band, valuation stops and a person decides. Outside the limits, valuation may record a rejection, and the more important behaviour is what it triggers next.

The usage decision is a hard stop in both industries. This is the moment the material becomes usable or does not. It moves stock out of quality inspection into unrestricted use, it writes a quality score, it can transfer inspection results into batch classification where they will feed every downstream certificate and batch determination, and it is one of the three points where SAP's own digital signature framework exists to capture who decided, when, and what the signature meant. That signature framework is not decoration. It is the system telling you, in its own architecture, that a person is expected here.

Yes, automatic usage decisions exist in standard SAP, restricted to skipped lots and lots where nothing was rejected and every required characteristic was confirmed. In a low-risk, non-critical material stream, that is a legitimate configuration. For an API, a sterile component, or a flight-critical forging, switching it on is an argument you will have with an auditor and lose.

Beyond the usage decision sit the decisions no automation should touch at all. Deciding that an out-of-specification result was a laboratory error is one. The FDA's guidance on investigating such results, revised in 2022, sets out a laboratory phase that looks for a documented, assignable cause, and the whole weight of that guidance rests on the principle that you cannot invalidate a result because you do not like it and you cannot test into compliance. An algorithm that quietly discards an outlier is doing the one thing the guidance was written to prevent.

Accepting a nonconformance is another. In aerospace, a part that does not meet the drawing is not made acceptable by anyone at the supplier, no matter how confident anyone is that the deviation is harmless. It requires a concession from the authority that owns the design. In pharma, certification of the batch by the Qualified Person is personal, and no system output substitutes for that act.

Diagram illustrating the transition point where an automated machine process hands over control to a human operator.

What changes in pharma, and what changes in aerospace

The stop points look similar in shape and differ in what triggers them.

In pharmaceutical quality management, the trigger is usually a result and a clock. A recurring inspection fires when a batch reaches its retest date, and a machine can absolutely manage the scheduling, the sampling, and the recording. What it cannot manage is the consequence of an unexpected result on a batch already in stock. An out-of-specification finding on a retest starts an investigation with defined phases, potentially reaches other batches made from the same starting material, and may end in a recall decision. The document handling around that investigation is enormously helped by automation. The conclusion is not.

Two further pharma-specific constraints shape the design. Data integrity expectations mean every automated action needs to be attributable, legible, contemporaneous, original and accurate, which in practice means the extraction step must keep the source document, the page, and ideally the coordinates of the value it read, not just the value. And electronic signature rules require the signed record to carry the printed name of the signer, the date and time, and the meaning of the signature. SAP's digital signature captures all three, but only if the time zone configuration is right, which is the kind of detail that decides audits.

In aerospace the trigger is more often a characteristic and a chain of custody. A first article inspection report is a claim that a process, on a specific machine, with specific tooling, made a part that matches the drawing. Automating the comparison of 214 reported characteristics against 214 drawing requirements is genuinely valuable work and removes a known source of human error. Approving the first article is a different act, because it authorises production, and it belongs to a qualified reviewer.

Certificate authenticity carries more weight here as well. A heat treatment certificate from a special process supplier is only meaningful if that supplier held the relevant approval on the date the work was performed. A document extraction system can check the certificate number, the process specification revision, the date, and the approval status held on file, and can flag an expired or missing approval instantly. That is a control most organisations perform by sampling, if at all, and it is one of the clearest cases where automation makes the quality system stronger rather than merely faster.

And when a part is out of drawing, the material review process begins. The machine can assemble the evidence, pull the history of the same characteristic across previous lots, and route the package. The disposition of use-as-is, rework, repair, or scrap is a judgment with engineering liability attached to it.

Writing the stop rules down so they survive an audit

A stop rule that lives in someone's head is not a control. Five artefacts turn a sensible design into something you can hand to an inspector.

Start with an intended use statement per document type and per characteristic, written before anything is built, naming what the model reads, what it may conclude, what it may never conclude, and what happens when it is unsure. This is the document Annex 22 asks for, and it is also the document that prevents scope creep six months later when someone asks why the system cannot just approve the easy ones.

Set thresholds per characteristic with a written rationale. The batch number threshold and the appearance threshold should not match, and the reason they do not match should be one sentence a reviewer can read and agree with.

Build a challenge set that contains your actual pain. Not a clean sample of typical certificates, but the scanned fax from the supplier who has not updated their template since 2011, the certificate with the specification column adjacent to the result, the one reporting in percent when your specification is in ppm, the one with a handwritten amendment and an initial beside it. Measure performance on that set, and measure it against how the manual process performs on the same set. The requirement is that the automated process is at least as good, and the only way to make that claim is to have measured both.

Define the undecided behaviour explicitly. What the reviewer sees, how the item is queued, who owns the queue, what the service level is, and what happens if nobody picks it up. An automation that stops correctly and then strands the work has moved the failure rather than removed it.

Put the model inside change control. This is where most pilots quietly fail their first audit. A new model version is a change. A supplier changing their certificate layout is a change to the input space. A new material with different characteristics is a change to the intended use. Each needs assessment, and the system needs to record which model version produced which value on which date, because in two years somebody will ask exactly that about batch 24-0817.

The point of stopping well

There is a version of this conversation where regulation is the obstacle and AI is the thing straining against it. That framing produces bad systems. The constraints in GMP and in aerospace quality standards are not arbitrary. They encode the specific ways that quality decisions have gone wrong in the past, usually expensively and sometimes tragically.

A better framing is that the regulations have already done the hardest part of the design work. They have marked, in considerable detail, which decisions require a human with a name and delegated authority. Those marks are the stop points. Everything upstream of them is fair game, and there is far more of it than most quality teams assume.

For the API batch, that means the certificate is read, the values are normalised and recorded with their sources, the specification comparison is done including the guard band, the supplier's certificate is checked against the purchase order and the material specification held in the system, and the whole package is presented to the quality engineer as a complete, evidenced, pre-checked lot with two characteristics flagged for her attention. Her decision takes four minutes instead of forty. It is still her decision.

For the forging, it means 214 characteristics are compared automatically, the three that sit inside the guard band are highlighted, the heat treatment approval is verified against the date of processing, and the reviewer spends their time on the engineering question rather than on the clerical one.

That is what good looks like. Not an empty quality department, but one where every hour of expert judgment is spent on something that actually needs judgment. The machine's job is to arrive at the boundary with everything in order, and then to stop.

For the team building this in SAP

The stop points above map onto specific touchpoints. Lot creation and specification assignment happen at goods receipt through the inspection type on the material master. Results land through results recording, either in the standard worklist or through the inspection data interface if a laboratory system or extraction service posts them, with the relevant remote function modules sitting in the QIRF function group and the corresponding standard interfaces available for recording results and setting usage decisions. Characteristic valuation behaviour is controlled by the valuation mode on the sampling procedure, which is where the mean-versus-individual-value question is settled, and custom valuation logic is a supported extension point through the function module behind each valuation rule. Specification limits and the wider plausibility limits are separate fields on the master inspection characteristic, and the plausibility band is a useful place to catch unit errors before they reach valuation.

The usage decision is where authority is enforced. Digital signatures can be required at results recording, at the usage decision, and at physical sample drawing, configured through the QM material authorisation group on the material master QM view and a signature strategy that can chain individual signatures in sequence. Automatic usage decisions are available for skipped lots and for lots with no rejections, controlled per inspection type with a waiting time, and that switch is the one to leave off for critical materials. Restricting who may make a usage decision uses the standard authorisation objects for inspection type and for changing stock posting fields, and restriction by individual usage decision code is not standard, so it needs a custom check in the usage decision check user exit. Transfer of results into batch classification happens at the usage decision through the link between master inspection characteristics and class characteristics, which is also what feeds outgoing certificates, so an incorrect value written there propagates further than people expect.

 

Artificio builds document intelligence for regulated SAP environments, with confidence scoring at characteristic level, guard-band routing, layout change detection, and a full audit trail from the source document through to the inspection lot. If you are mapping your own stop points, we are happy to compare notes.

Lal Singh, SAP AI Automation Expert

CEO & Founder of Artificio

See it in your SAP environment

Request a demo

Bring us a document, a process, or a bottleneck. We'll show how Artificio captures, validates, and posts into SAP — then scale from there.

Request a demo

Security & compliance

Enterprise security across every solution

ISO 27001:2013 certified, SOC 2 Type 2 compliant, GDPR and HIPAA ready. Every agent action is logged, auditable, and runs in isolated environments.

  • ISO 27001:2013
  • SOC 2 Type II
  • GDPR ready
  • HIPAA ready