Skip to main content
All articles

Blog

Right Model, Right Job: Why Your Document Pipeline Should Never Use One Brain for Everything

Struggling with thousands of invoices every week? Learn how routing documents across fast AI, strong reasoning models, and human judgment cuts processing time by 80%.

Lakshay Sharma
Artificio's Document Pipeline Intelligence

It is 8:40 on a Monday morning, and Priya already knows what her week looks like. Over the weekend, 2,400 supplier invoices landed in the shared accounts payable inbox. Some are crisp PDFs from large vendors. Some are photos taken on a loading dock, slightly tilted, with a thumb in the corner. One is a fax. Yes, a fax, in 2026.

Priya runs accounts payable at a mid-size distributor, and her team has five people. Every one of those invoices has to be read, matched against a purchase order, checked against what actually arrived at the warehouse, approved, and queued for payment by Friday. Miss that window and suppliers start calling. Early-payment discounts quietly expire. The finance director starts asking why cash flow looks strange.

Here is the math that keeps her up at night. If a person spends four minutes on each invoice, the pile takes about 160 hours. Her team has roughly 200 working hours in the whole week, and those hours also have to cover vendor questions, month-end prep, and the audit request that arrived on Thursday. The pile does not fit. It never fits.

This is the moment where most companies reach for one of two ideas, and both of them are wrong in interesting ways. This post is about the third idea, the one that actually works. It comes down to five words. Right model, right job.

The Monday Morning Problem

Let us start with the two bad ideas, because they are both tempting.

The first idea is to hire more people. It works, in the sense that more hands clear more paper. But the cost grows in a straight line with the volume. Double the invoices and you double the team. Worse, most of what those extra people do is boring. They read the same kind of invoice from the same kind of vendor and confirm that the numbers add up. Nobody went into finance dreaming of confirming that 14 plus 22 equals 36 for eight hours a day.

The second idea sounds more modern. Buy the most powerful AI model available and point it at the whole pile. Every invoice goes through the smartest brain in the building. And here is the surprise. This often works too, at least in the sense that the answers come out accurate. The trouble shows up on the bill.

The most capable AI models cost far more to run than the small, fast ones. One engineer who tracks these prices estimated that the gap between a cheap model and a frontier model in 2026 is roughly 100 times per token. Most documents never needed that kind of power. A clean invoice from a vendor you have paid 40 times before is not a philosophical puzzle. It is a form. Sending it to the most expensive reader available is like hiring a structural engineer to tell you whether a door is open or closed.

Think of a hospital emergency department. If the chief surgeon personally greeted every patient who walked in, checked every sore throat, and took every blood pressure, the surgeon would be exhausted by noon and the person with the real emergency would wait. Hospitals solved this long ago with triage. A fast, simple check at the front door decides who needs a nurse, who needs a doctor, and who needs the surgeon. Each patient gets the right level of care, and the surgeon spends the day doing surgery.

Documents deserve the same system.

Three Seats at the Table

At Artificio, we think about document processing as a team with three seats, and every document is assigned to exactly one of them.

The first seat belongs to a fast, inexpensive AI model. Its job is routing and routine checks. It opens every single document, figures out what it is, pulls out the important details, and runs the basic sanity checks. Is this an invoice or a credit note? Does the vendor exist in our records? Do the line items add up to the total? Has this exact invoice been submitted before? When the answer to all of those is a clean yes, the document keeps moving with no further attention.

The second seat belongs to a stronger, more expensive AI model. It does not see everything. It only sees the hard matches, the documents where the first seat noticed something that did not line up neatly. The invoice says "12 pallets of stretch film" and the purchase order says "film, stretch, 500mm, 12 units." Same thing? Probably. But somebody has to reason about it, and that somebody needs more horsepower than the front door has.

The third seat belongs to people. Not as a fallback for broken technology, but as the right place for decisions that involve judgment, relationships, policy, and risk. A vendor changed their bank account number the day before a large payment. A supplier is disputing a credit. An invoice is 18 percent higher than the contract allows. These are not reading problems. They are business decisions, and people should make them.

So the full idea fits in two sentences. A cheap model handles routing and routine checks. A stronger one takes the hard matches, and people handle the rest.

An empty meeting table with three seats surrounding a single central pile of paper documents. 
 

Notice what this design does not do. It does not try to make one model do everything, and it does not treat people as a failure state. Each seat exists because it is genuinely the best place for a certain kind of work. That is the whole philosophy, and the rest of this post is about how it plays out in practice.

Seat One: The Fast Reader

Let us follow a single invoice through the system. It comes from Harlow Packaging, invoice number INV-88213, for $4,812.50. It arrived as a PDF attachment at 6:12 on Sunday morning, which tells you something about how small suppliers run their weeks.

The fast reader opens it first. In about the time it takes you to blink, it recognizes this as an invoice rather than a statement or a delivery note. It reads the vendor name, the invoice number, the date, the purchase order reference, and each line item. Then it runs a short list of checks. Does Harlow Packaging exist in the vendor list? Yes. Is INV-88213 already in the system? No. Do the three line items add up to $4,812.50? They do. Does the purchase order number on the invoice, PO 55017, exist and still have an open balance? Yes.

Everything lines up, so the invoice moves straight through. No stronger model gets involved. No person looks at it. The whole trip takes a few seconds and costs a tiny fraction of a cent.

This is the quiet heart of the system, and it deserves more respect than it usually gets. In most businesses, the bulk of documents are ordinary. They look like the ones that came before them. Industry write-ups on model routing consistently report that most production traffic is easy, with one practitioner putting the share of requests that never needed the top-tier model at around 80 percent. Your own numbers will differ, but the pattern is familiar to anyone who has watched an inbox for a week. Most of it is routine.

The fast reader has a second job that matters just as much. It routes. When something looks off, the fast reader does not try to be a hero. It does not guess. It writes down what it noticed and passes the document to the next seat. A handwritten delivery note with a smudged quantity gets flagged. An invoice that mentions two purchase orders gets flagged. A line item with no obvious match gets flagged.

This is the opposite of what people expect from AI. We tend to imagine the cheap model as the sloppy one. In a well-designed system, the cheap model is the honest one. Its most valuable skill is knowing the limits of its own judgment and saying so quickly. A front-desk clerk who admits "this one is above my level" saves everyone more time than one who bluffs.

Seat Two: The Specialist

Now picture the other invoice from Monday's pile. This one comes from a vendor called Brightline Industrial, for $17,940.00. The purchase order, PO 55102, lists the goods as "hydraulic hose assembly, 3/4 inch, 50 units." The invoice lists them as "HH-34 hose kit x 5 cartons." Meanwhile, the warehouse receiving record says 40 units arrived on Thursday, with the remaining 10 backordered.

A fast reader sees the mismatch immediately and flags it. It cannot sensibly resolve it, because resolving it requires reasoning. Is a carton of HH-34 hose kits ten assemblies? The vendor catalog says yes. Does that make five cartons equal to 50 units, which matches the order? Yes. But only 40 units arrived, so should the invoice be paid in full, or only for what showed up?

This is where the stronger model earns its price. It reads the invoice, the purchase order, and the receiving record together. It checks the vendor catalog to confirm what a carton means. It concludes that the invoice bills for 50 units while only 40 were received, and that the correct action is to match 40 units now, hold the remaining amount, and note the backorder. It writes down its reasoning so that anyone can review it later.

Matching problems like this one are where cheaper models tend to run out of road. The work is not about reading text. It is about holding several documents in mind at once, understanding that different words can describe the same thing, and noticing when the numbers tell a different story than the descriptions. A stronger model is simply better at that kind of reasoning, which is why it belongs in the second seat.

Here is the part that matters for the budget. The stronger model only sees a slice of the pile. If the fast reader sends along 25 percent of documents, the expensive model does 25 percent of the work, not 100 percent. The price of its power is paid only where power is needed.

Researchers have been testing this exact design. A June 2026 paper on cascaded model serving described a two-stage system that sends easy requests to cost-effective models and escalates any answer judged low quality to a stronger one. It reported keeping 97 to 99 percent of the strongest model's accuracy. Another recent study of cascades with cost limits found that a well-tuned setup could reach close to top accuracy while spending less than a fifth of what the most powerful model would cost on its own. The numbers vary with the task, but the direction is consistent. Mixed teams of models do almost as well as the best model alone, at a fraction of the price.

Seat Three: The People

Now we come to the seat that gets misunderstood the most.

In a lot of AI conversations, humans show up as a safety net, the thing you keep around until the technology gets good enough to remove them. That framing misses what people actually contribute. Some decisions are not hard because the information is difficult to read. They are hard because they involve judgment about risk, relationships, or rules that live in people's heads.

Back to Monday. A third invoice arrives from Delmar Logistics for $9,260.00. Everything on it looks fine, with one exception. The bank account listed for payment is different from the one Delmar has used for the last three years.

The fast reader catches it. The stronger model confirms it is unusual. And now the system does something very deliberate. It stops. It does not try to decide whether this is a legitimate banking change or a fraud attempt, because that is not a reading question. Changed bank details are one of the oldest tricks in the payment fraud playbook, and the right response is a phone call to a known contact at Delmar, made by a person who knows the relationship. The system places the invoice in a review queue, shows the reviewer exactly what changed and why it was flagged, and waits.

Other items land here too. A supplier credit that does not match any open invoice. A price 18 percent above the contract rate. A new vendor that needs to be set up. A very large payment above the approval limit. None of these are failures of AI. They are the cases where a human decision is the product.

Good design also makes this seat efficient. The reviewer does not receive a raw document and a blank screen. They see the original invoice beside the extracted data, with the uncertain fields highlighted and a short note explaining what triggered the flag. Industry guidance on human-in-the-loop document processing describes this as exception handling rather than routine processing, and one widely cited guide estimates that surfacing only the low-confidence fields and edge cases cuts manual review effort by roughly 70 percent in many rollouts. Priya's team does not read 2,400 invoices. They review the few that genuinely need a person.

And there is a bonus. Every correction a person makes becomes a lesson. When a reviewer confirms that "HH-34 kit" means ten units per carton, the system remembers. Next time, that match belongs to the specialist, or even to the fast reader. The people seat slowly shrinks the amount of work that needs the people seat. That is the opposite of a safety net. It is a teacher.

How a Document Decides Where to Sit

A fair question at this point is how the system knows which seat a document belongs in. Nobody is standing at a desk sorting paper. The answer is something called confidence.

Every time the fast reader extracts a value, it also produces a measure of how sure it is. Think of it as a dial from zero to one hundred. A clearly printed invoice number on a clean PDF might come back at 99. A smudged quantity on a photographed delivery note might come back at 62. The system sets thresholds, and those thresholds decide the route. Above the line, the document keeps going. Below it, the document moves up a seat.

Teams that run these systems often start with a line somewhere in the range of 90 to 95 percent for money fields, and then adjust field by field. The total on an invoice deserves a higher bar than the vendor's mailing address, because a wrong total costs real money and a wrong address usually does not.

Two other signals matter just as much. The first is rules. Even a document with perfect confidence gets escalated if it breaks a business rule, like an amount above the approval limit or a purchase order that has already been fully billed. The second is history. A familiar vendor sending a familiar kind of invoice earns a smoother ride than a stranger sending something unusual.

There is one more wrinkle, and it is worth understanding because it shows how much careful thought sits behind a design that looks simple. A May 2026 research paper pointed out that a pure cascade, where the cheap model always goes first, has a built-in cost. Every document pays the cheap model's price before any decision is made, even the ones that were always going to need the stronger model. That paper found that a lightweight router that decides upfront, before any model reads the document, beat the best cascade on four of five datasets.

What does that mean in practice? It means the smartest systems do not treat the front door as a single step. Some documents are obviously hard before anyone reads them closely. A handwritten delivery note in poor light, a 40-page contract with scanned amendments, an invoice in a layout the system has never seen. A quick look at the file itself is enough to send those straight to the specialist, skipping the wasted first attempt. Easy documents race through. Hard ones skip the line. Everything else gets the standard route.

None of this is magic. It is the same instinct a good dispatcher has when they glance at a call and know exactly who to send.

The Monday Pile, Re-Run

So what happens to Priya's 2,400 invoices when the three seats are working together?

The fast reader opens all 2,400. About 1,800 pass every check cleanly and flow straight through to approval. That is 75 percent of the pile with no human involvement and no expensive model.

The other 600 get flagged. The stronger model takes those, reads them alongside the purchase orders and receiving records, and resolves about 480 of them. These are the carton-versus-unit puzzles, the partial shipments, the description mismatches that look different but mean the same thing.

That leaves 120 invoices for people. The changed bank detail. The disputed credit. The 18 percent price jump. A reviewer looks at each one for around 10 minutes, since the system has already done the digging and written up what it found.

Now compare the effort. In the all-people version, 2,400 invoices at four minutes each is 160 hours. In the routed version, 120 invoices at ten minutes each is 20 hours. Priya's team gets almost an entire week back. Not by working faster, but by working only on the things that truly need them.

The money story is just as clear, using a simple cost scale. Say the fast reader costs 1 unit per document and the stronger model costs 20 units, which is a modest gap compared with the 100 times difference some engineers report. Running all 2,400 invoices through the stronger model costs 48,000 units. In the routed system, the fast reader handles all 2,400 for 2,400 units, and the stronger model handles 600 for 12,000 units. The total is 14,400 units. That is 30 percent of the original cost, a saving of 70 percent, while every document still gets the right level of attention.

These figures are an illustration, not a promise, and your own mix of documents will change them. A pile full of messy handwritten forms will send more work to the upper seats. A pile of standard vendor invoices will send less. But the shape of the saving stays the same. Engineering write-ups on deployed routing systems commonly describe cost reductions of two to four times at equivalent quality, and narrow, well-understood workloads can go much further.

A tall stack of documents and paperwork labeled

Where This Goes Wrong

Any honest description of this approach has to include the ways it can fail. Routing is not a free lunch, and the failures are worth knowing because they are easy to miss.

The biggest risk is a judge that is too generous. If the fast reader is too eager to pass documents through, errors slide past quietly. Everything looks great on the cost dashboard while the quality of the output slowly slips. One practitioner put it bluntly. If routing is wrong, hard requests land on the cheap model, the product gets worse, and every cost chart turns green. The defense is measurement. Sample the documents that passed automatically and have people check them regularly. If the error rate creeps up, tighten the thresholds.

The second risk is the opposite problem. A judge that is too nervous sends everything upward and the savings evaporate. The system works, but it works expensively. Watching the share of documents at each seat tells you quickly whether the dial is set sensibly.

The third risk is the slow decay of rules. Vendors change their invoice layouts. Companies add new approval policies. A system calibrated in January may be behaving oddly by June. That is why the correction loop matters so much. When reviewers fix mistakes, those fixes should flow back into the system. Teams that treat exception handling as the product, rather than an afterthought, tend to see durable results.

The fourth risk is human, not technical. If the review queue is overloaded or poorly designed, reviewers start rubber-stamping. A person who clicks approve on the 80th item in a row is not providing oversight. The fix is to keep the queue small, show reviewers exactly what matters, and make the hard cases feel like meaningful work rather than a pile of leftovers.

None of these problems are reasons to avoid the design. They are reasons to build it carefully, with monitoring from day one rather than after the first surprise.

Beyond the Invoice

Accounts payable makes a clean example, but the same three seats show up wherever documents pile up.

Think about an HR team onboarding 300 new hires in a hiring season. Most of the paperwork is straightforward, like signed offer letters and standard identification forms. The fast reader handles those. A small share involves unusual details, like a name that differs between documents or a start date that conflicts with a contract. The stronger model sorts out most of those. A few need a person, like a credential that cannot be verified. Same shape, different paper.

Or think about a university admissions office during application season, with thousands of transcripts arriving in dozens of formats from schools around the world. Most are readable and routine. Some use unfamiliar grading scales that need careful interpretation. A handful raise questions only an admissions officer can answer.

The same is true for insurance claims, loan files, shipping paperwork, and customs entries. The logistics industry has learned this lesson the hard way. Systems built on clean-document accuracy look great in demos and stumble at the exception queue, where mismatched bills of lading and ambiguous customs entries either get caught or get posted. The teams that succeed treat the exception queue as the center of the design, and tune confidence thresholds for each field before launch.

Wherever you see a large volume of mostly routine paper with a small tail of tricky cases, you are looking at a place where the right model for the right job pays off.

Choosing Your Own Seats

If you are thinking about how this might apply to your own operation, a few questions are worth asking before anything else.

What share of your documents are truly routine? Pull a sample of a few hundred and sort them by feel. The honest answer is often higher than people expect, and it tells you how much work belongs in the first seat.

Which mistakes are expensive? A wrong mailing address is an annoyance. A wrong payment amount or a payment sent to the wrong account is a real loss. Those fields deserve higher confidence bars and quicker escalation.

Where do your people spend their time today? If your best reviewers spend most of the day confirming that totals add up, that is the work to hand to the first seat, so their judgment can go where it matters.

And how will you know if it is working? Decide in advance how you will sample automated results, how you will track the share of documents at each seat, and how corrections will flow back into the system.

These are not technical questions. They are operational ones, and the answers shape the design more than any choice of model.

What Priya Does With Her Friday

Let us return to Priya, because the point of all this is what happens to the people involved.

By Wednesday afternoon, the Monday pile is done. The 1,800 straightforward invoices were approved and queued before she finished her first coffee. The stronger model resolved the carton puzzles and partial shipments overnight, with notes attached. Her team worked through the 120 flagged items in two days, and the phone call to Delmar Logistics confirmed that the bank change was fake. A fraudulent payment of $9,260.00 never left the building.

On Thursday, she has time for the audit request. On Friday, she has time to call two vendors back, and to sit down with her team and talk about what they want to work on next quarter. Nobody stayed late. Nobody read 2,400 invoices.

That is what right model, right job really means. It is not about replacing people, and it is not about worshipping the most powerful technology. It is about matching each piece of work with the part of the system best suited to do it. The fast reader keeps the line moving. The specialist untangles the hard matches. And people do what only people can do, which is make the calls that carry real consequences.

At Artificio, this is how we think about every document that enters the system. The goal is never to run the biggest model on the most paper. The goal is to get each document to the right seat, quickly and at the right cost, so the people on your team can spend their week on work that deserves them.

Your pile is waiting. The question is who, or what, should read it first.

Lakshay Sharma

Project Manager

See it in your SAP environment

Request a demo

Bring us a document, a process, or a bottleneck. We'll show how Artificio captures, validates, and posts into SAP — then scale from there.

Request a demo

Security & compliance

Enterprise security across every solution

ISO 27001:2013 certified, SOC 2 Type 2 compliant, GDPR and HIPAA ready. Every agent action is logged, auditable, and runs in isolated environments.

  • ISO 27001:2013
  • SOC 2 Type II
  • GDPR ready
  • HIPAA ready