Picture a Tuesday morning in the accounts payable team of a mid-sized manufacturer outside Stuttgart. A supplier calls to ask why invoice 4471-B, worth โฌ18,420 in hydraulic fittings, was paid at a different unit price than the one on the purchase order. The AP lead opens SAP and finds the posting. It looks clean. The document number is there, the GL account is right, the vendor is right. The posting user is a technical account called AGENT_AP_01.
So she asks the obvious question. What happened here?
The automation vendor's dashboard says the invoice was "processed successfully" at 2:14 a.m. That is all it says. Nobody can tell her which price the agent read, whether it compared the invoice against the purchase order or against an older contract, or why it decided the difference was acceptable. The supplier is waiting on the line. The quarterly audit is three weeks away. The only honest answer anyone can give her is "we do not know."
This scene plays out in more SAP shops than most vendors would like to admit. For a few years, the pitch for AI in SAP was about speed. Fewer clicks, faster postings, smaller backlogs. In 2026, SAP changed the conversation. Through a new API policy, a new gateway for agent traffic, an agent registry and an overhauled partner certification program, SAP has made it clear that any agent touching its systems needs to be identifiable, governed and traceable. Its own agents are being built that way. Every third-party agent will now be measured against that standard, whether the vendor likes it or not.
This post gives SAP customers a practical way to do that measuring. It comes down to three questions you should be able to answer for every action any agent takes. What did it do? Why did it act? Which data did it use? If a vendor cannot answer all three for a single document, on demand, their agent is not ready for the environment SAP is building.
What SAP Actually Changed This Year
A lot of the coverage of SAP's 2026 moves has focused on commercial fights, and those fights are real. Before getting to the checklist, here is what changed, stated plainly.
The first change came through SAP's API policy. The April 2026 version limits the use of SAP APIs for interaction with semi-autonomous or generative AI systems that plan, select or execute sequences of API calls. Read literally, a third-party agent should not be deciding on its own how to fetch and move data inside your SAP system. SAP CEO Christian Klein said on the company's first-quarter call that the goal is to protect SAP's domain knowledge and system performance, not to block customers from their own data. User groups, including the German-speaking user group DSAG, have pushed back and asked for clearer terms.
The second change is the route agent traffic is supposed to take. SAP has pointed to Joule, Business Data Cloud and its Agent Gateway as the approved paths for AI to work with SAP data. The gateway sits in SAP Integration Suite. It exposes curated SAP APIs as managed MCP servers and adds metering, rate limiting, token tracking and agent identity verification. Third-party agents are expected to reach SAP processes by talking to Joule agents through the Agent-to-Agent protocol, usually shortened to A2A.
The third change is enforcement. SAP tied its data extraction rules to a security patch that began blocking noncompliant ODP-RFC calls from June 9, 2026, and published a self-assessment note so customers can audit their existing usage.
The fourth change is governance tooling. At Sapphire in May 2026, SAP opened up its AI Agent Hub to a much wider set of customers. The hub includes a registry that discovers agents, models and MCP servers regardless of who built them. It adds a unique identity for each agent through SAP Cloud Identity Services, session-level observability that tracks things like whether an agent called the right tools, and process mining through SAP Signavio to check whether an agent followed its intended path. SAP scheduled several of these capabilities for the third quarter of 2026. Joule Studio 2.0 added logging that SAP describes as SOX-compliant, along with human approval steps inside agent flows.
The fifth change is certification. SAP's Integration Certification Center split its partner program into two tiers. The premium tier, Integration Certification, is for solutions aligned with SAP BTP and clean core. A lighter Interoperability Review covers solutions that are technically sound but fall short of the strategic criteria. The new framework was set to launch in the third quarter of 2026.
You can argue about whether SAP is opening its platform or closing it. Analysts at Forrester have been openly critical, warning that the API rules could revive the anxiety of the 2017 indirect access disputes. That debate matters for your contract negotiations. It does not change the practical reality for your operations team. SAP is building an environment where every agent has a name, a permission set and a record of its sessions. A third-party agent that shows up with nothing more than a success counter will look badly out of place.
Why "It Worked" Is No Longer Enough
Most automation buyers in the SAP world evaluated tools on outcomes. Straight-through processing rate. Accuracy on a test batch. Hours saved per month. Those numbers still matter. But they describe what happened on average, across thousands of documents. Auditors, controllers and regulators do not ask about averages. They ask about one document.
Think about how a good accounts payable clerk works. If you ask why they approved an invoice with a 3% price variance, they can tell you. They looked at the purchase order, saw the contract allowed a fuel surcharge, checked the surcharge schedule for March, and confirmed the amount matched. They can point to the page they read. That ability to explain one decision, from start to finish, is what makes a person trustworthy in a finance role.
An agent needs to meet the same standard. In some ways it needs to beat it, because an agent can make the same mistake ten thousand times overnight before anyone notices. When something goes wrong in a human process, you talk to the person. When something goes wrong with an agent, the record it left behind is all you have.
There is also a structural shift underway. In the old model, a third-party tool pulled data out of a document, handed it to SAP through an interface, and SAP did the posting. The tool's job ended at the edge of SAP. Agents blur that edge. They read a document, decide what it means, choose which transaction to run, and sometimes trigger follow-up actions like a payment block or an email to the supplier. Every one of those decisions is a point where the business needs evidence. SAP's new rules are, among other things, an attempt to make sure those decisions run through infrastructure that can be seen and governed.ย
So the question for any SAP customer is simple. If your vendor's agent made a bad call at 2:14 a.m., could you reconstruct exactly what happened by 9:00 a.m.?
The Three Questions Every Agent Must Answer
Every capability that matters for agent accountability fits into three questions. They are easy to remember and hard to fake.
What did it do? This is the action record. Every step the agent took, in order, with timestamps, the system it touched and the result it got back.
Why did it act? This is the reasoning record. What triggered the agent, which rule or policy it applied, how confident it was, and what alternatives it rejected.
Which data did it use? This is the evidence record. The exact source document, page and field, the version of the master data it read, and anything it pulled from outside the document.
Sitting underneath all three is a quieter question. Who allowed it? An agent should act under its own identity, with permissions that someone in your company approved, and it should stop and ask a human when it reaches the edge of those permissions. SAP's move toward unique agent identities makes this foundation much easier to inspect.

Here is the checklist in detail. Use it in vendor evaluations, in renewal conversations and in internal reviews of the agents you already run.
Question One: What Did It Do?
The action record sounds basic, and most vendors will tell you they have it. Push on the details. A log line that says "Invoice processed" is not an action record. It is a receipt for a black box.
A real action record for one invoice should read like a sequence of moves you could replay. The agent received the email at 01:52. It classified the attachment as a vendor invoice. It extracted 23 fields. It looked up vendor 100482 in the vendor master. It matched the invoice to purchase order 4500019873. It ran a three-way match against goods receipt 5000044120. It found a price variance on line 2. It applied the tolerance rule. It created the posting through an approved interface and received document number 5105600231 back. Each of those lines carries a timestamp and a status.
Use this checklist for the action record.
- โ Every step is logged individually, not just the final outcome.
- โ Each step records the system touched, the operation called and the response received, including SAP document numbers.
- โ Failed and retried steps are logged, not just the successful ones.
- โ The agent acts under its own named identity, never a shared or human user account.
- โ Logs are tamper-evident and kept for at least as long as your financial record retention policy requires.
- โ You can export the full record for one document in a format your auditors can read without the vendor's software.
- โ Follow-up actions such as payment blocks, supplier emails and workflow tasks appear in the same record as the main action.
The identity point deserves extra attention. Many older integrations post into SAP using a generic technical user that several tools share. If three automation products all post as BATCH_USER, nobody can tell which one created a bad entry. SAP's direction, with agent identities managed through Cloud Identity Services and verified at the gateway, makes shared accounts look like what they are. A gap.
Question Two: Why Did It Act?
This is where most third-party agents fall apart. Plenty of tools can tell you what they did. Very few can tell you why.
"Why" has several layers. The first is the trigger. Did the agent start because an email arrived, because a scheduled job ran, or because another agent asked it to? With A2A making agent-to-agent calls more common, a chain of agents calling each other can become very hard to untangle unless each one records who asked it to act.
The second layer is the rule. When the agent accepted a 2.8% price variance on line 2, which tolerance did it apply? Was it your company-wide 3% tolerance, a vendor-specific exception, or a number the vendor configured during onboarding and never told you about? The record should name the policy and its version.
The third layer is confidence. A good agent knows how sure it is. Reading the unit price as โฌ42.60 with 99% confidence is a different situation from reading it as โฌ42.60 with 71% confidence because the scan was blurry. The confidence score belongs with the decision, and you should be able to see the threshold that sends a document to a person instead.
The last layer is the paths not taken. If the agent considered two possible purchase orders and picked one, the record should show both and explain the choice. A large share of real-world errors come from an agent confidently picking the wrong candidate when two looked alike.
Use this checklist for the reasoning record.
- โ The trigger for every action is recorded, including which user, schedule or agent started it.
- โ The business rule or policy applied is named, along with its version.
- โ Confidence scores are stored per field and per decision, not just per document.
- โ Thresholds that route work to a human are visible to you and configurable by you.
- โ When the agent chose between candidates, such as two purchase orders or two GL accounts, the rejected options and the reason for rejecting them are recorded.
- โ Human approvals and overrides are captured with the approver's name, time and comment.
- โ The explanation is written in business language a controller can read, not as raw model output.
That last item matters more than it looks. Some vendors respond to "why" by dumping a long block of model reasoning into a log file. That is not an explanation. A controller needs one or two sentences that tie the decision to a policy and a piece of evidence. "Accepted a 2.8% variance on line 2 under tolerance policy AP-TOL-03, version 4, which allows up to 3% for vendor class B" is an explanation. Ten paragraphs of model output is homework.
Question Three: Which Data Did It Use?
The evidence record is the one that saves you during an audit. It answers the question every auditor eventually asks. Show me where that number came from.
For document work, the answer should point to a specific place. The unit price of โฌ42.60 came from page 1, line 2 of the PDF, in the Unit Price column. The payment terms of Net 45 came from the footer on page 2. A strong agent can show you the original document with that region highlighted, so a reviewer can check the agent's reading against the source in seconds.
Documents are only half of it. Agents also read SAP data, and that data changes. The agent matched the invoice against purchase order 4500019873, but which version of it? If someone changed the price on the purchase order on Wednesday and the agent processed the invoice on Tuesday night, you need to know the agent saw the old price. The same goes for vendor master data, bank details and tolerance settings. An evidence record should capture what the agent saw at the moment it acted, not what the data looks like today.
A third layer has become urgent this year. Where did the agent get SAP data from, and through which channel? Under SAP's 2026 API policy, the path matters. An agent that reads data through an approved interface leaves a trail SAP's own tools can see. An agent that reaches in through an old extraction method could be out of policy, and it may simply stop working when a patch closes that door. Your vendor should be able to tell you, in writing, exactly which interfaces their agent uses to read and write SAP data.
Use this checklist for the evidence record.
- โ Every extracted value links to its exact page, region and field in the source document.
- โ Reviewers can see the source document with the relevant region highlighted.
- โ The record captures the version of SAP master and transactional data the agent read at decision time.
- โ Any outside data, such as exchange rates, tax tables or supplier portal entries, is recorded with its source and timestamp.
- โ The vendor documents every SAP interface the agent uses, and each one sits on a path SAP supports for AI access.
- โ Data the agent stores outside SAP, like copies of invoices or extracted fields, has a defined location, retention period and deletion process.
- โ Personal data in documents is identified and handled according to your privacy obligations.
Putting the Checklist to Work With the One-Document Test
Checklists are easy to nod along to. Vendors will say yes to every line in a sales meeting. The way past that is a simple test you can run in any evaluation or renewal. Pick one real document from last month and ask the vendor to show you its full story.
Choose something a little messy. An invoice with a price variance, a partial delivery or a handwritten note works well. Give the vendor the document number and ask them to walk you through, live, what their agent did with it, why it made each decision, and which data it relied on. Do not accept slides. Ask for the actual record.
Here is what a strong answer looks like, using the same โฌ18,420 invoice from the opening.
The email arrived at 01:52, and the agent logged the trigger as a mailbox rule on the AP inbox. It classified the PDF as a vendor invoice with 98% confidence. It extracted the vendor name, invoice number, purchase order reference, three line items, VAT and totals, each linked to a highlighted region on the page. It looked up the vendor in SAP and recorded the version of the vendor master it read. It matched the invoice to the referenced purchase order and goods receipt. It found that line 2 was billed at โฌ42.60 per unit against a purchase order price of โฌ41.45, a 2.8% difference. It applied the company's 3% tolerance policy for that vendor class and recorded the policy name and version. Because the variance sat within tolerance but above 2%, the agent routed the invoice to a human approver, exactly as the policy required. The AP lead approved it at 08:31 with a note that the supplier had announced a surcharge. The agent then posted the invoice through the approved interface and stored the SAP document number it got back.
Every one of those sentences maps to a line in the record. That is what ready looks like.

Now compare that with a weak answer. The vendor shows a dashboard. It says the invoice was "auto-processed" with "high confidence." When you ask about the price variance, they say the agent "handles tolerances automatically." When you ask which purchase order version it used, they need to check with engineering. When you ask how it writes to SAP, the answer involves an RFC connection someone set up years ago. None of this means the vendor is dishonest. It means their product was built for a world where nobody asked these questions. That world is ending.
Red Flags to Watch For
A few patterns tend to show up in agents that will struggle under SAP's new expectations. If you see any of these, dig deeper before signing or renewing.
The agent posts under a shared user account. If the vendor cannot give their agent its own identity in your landscape, you lose the ability to separate its actions from everything else in the system.
The vendor cannot say which SAP interfaces the agent uses. Vague answers like "we connect through standard APIs" are not enough in 2026. Ask for the list.
Explanations only exist inside the vendor's interface. If you cannot export the record for one document and hand it to an auditor, the evidence belongs to the vendor and not to you.
Confidence is a single number for the whole document. A document-level score hides the one field that was wrong. Ask for field-level confidence.
There is no human checkpoint by design. Good agents know when to stop. If a vendor brags that nothing ever goes to a person, ask what happens when the agent is unsure.
The vendor treats SAP's new rules as someone else's problem. A partner who has not read the 2026 API policy, has no plan for the Agent Gateway or A2A, and is not tracking SAP's certification tiers is betting that the rules will not apply to them. You are the one carrying that bet.
Eight Questions for Your Next Vendor Meeting
If you only have thirty minutes with a vendor, these eight questions will tell you most of what you need to know.
- Can you show me the complete record for one document I choose, right now?
- Under what identity does your agent act in our SAP system, and who approves its permissions?
- Which SAP interfaces does your agent use to read and write data, and how do they fit SAP's 2026 API policy?
- Where do you stand with SAP's Integration Certification or Interoperability Review?
- How does your agent record which version of our master data it used?
- Can I see field-level confidence scores, and can I change the thresholds that send work to a person?
- How long do you keep logs, where are they stored, and can we export them in an open format?
- If your agent and a Joule agent both touch the same process, how will we see the full chain of actions?
Write down the answers. Better still, ask for them in writing. When SAP's governance tools are fully in place and your auditors start asking about agents by name, you will be glad you did.
How Artificio Approaches Agent Accountability
At Artificio, we build AI agents that read and act on business documents, and many of our customers run SAP. We have been answering the "what, why and which data" questions for a long time, because document work demands it. When an agent reads a bill of lading, a certificate of analysis or a supplier invoice, the business needs to know exactly where every value came from.
That is why our agents link every extracted value back to its place in the source document, so a reviewer can check the reading in seconds. Confidence is tracked at the field level, and teams set the thresholds that send a document to a person. Every decision is logged with the rule that drove it, and human approvals are captured in the same record as the agent's own actions. The result is a story for each document that a controller can read and an auditor can test.
We also see SAP's 2026 changes as a healthy correction for the market. Customers should expect agents to be identifiable, governed and explainable, no matter who built them. Vendors who already work this way have nothing to fear from a higher bar. The ones who do not will need to catch up quickly.
If you want to run the one-document test on your own documents, we would be glad to show you how our agents handle it. Bring your messiest invoice.
The Bar Moved. Your Checklist Should Too.
SAP did not invent the idea that automation should be accountable. Auditors, controllers and regulators have expected it for decades. What SAP did in 2026 was build accountability into the platform itself, with agent identities, a gateway that knows who is calling, a registry that sees every agent, and observability that watches sessions in detail. That makes the gap between well-governed agents and opaque ones visible in a way it never was before.
For SAP customers, this is an opportunity as much as a burden. You now have a clear standard to hold every vendor to. Ask what the agent did. Ask why it acted. Ask which data it used. Ask for proof on one real document. The vendors who can answer quickly are the ones you can trust with your close, your payables and your audit.
Think back to the AP lead from the opening, staring at a posting by AGENT_AP_01 with no way to explain it. With the right agent, that supplier call takes two minutes. She opens the record, sees the surcharge, sees the approval, and sends the supplier a screenshot of the highlighted line. Then she gets on with her day.
That is the bar. Make sure your automation clears it.