TL;DR
- Sort failures by whether retrying could ever work. Everything else in the design follows from that classification.
- Transient failures retry with backoff and a ceiling. Permanent failures stop immediately and route to a person.
- Ambiguous failures retry with the same idempotency key, which is what makes the retry safe.
- A dead letter queue needs a named owner, depth monitoring, and a path back with history intact.
- See how Merlin Intake handles procurement requests. Request a demo.
Sort failures by whether retrying could ever work. Everything else in an error-handling design follows from that one classification, and most integrations never make it explicitly, which is why their behavior surprises everyone later.
An intake-to-ERP integration produces three kinds of failure: transient ones that succeed on retry, permanent ones that never will, and ambiguous ones where nobody knows whether the action happened. Treating all three the same way produces both infinite retry loops and requests that vanish without trace, which are the two complaints most often raised about integrations that were signed off as working.
This is written for procurement operations leaders specifying an integration with IT, or diagnosing one already running in production and behaving unpredictably.
What are the three kinds of failure?
- Transient. A timeout, a rate limit, a service restarting. The request was well formed and the conditions were temporarily wrong. Retrying is correct and will usually succeed within a few attempts.
- Permanent. A rejected cost center, a missing required field, a supplier that does not exist in the ERP. These surface later as match exceptions. The request will fail identically on every attempt because the problem is in the request rather than the conditions. Retrying is not just useless, it generates thousands of identical error records that bury the ones worth reading, which is how a monitored integration becomes an unmonitored one without anybody deciding to stop watching.
- Ambiguous. No response arrived. The action may have succeeded, may have failed, and the sender cannot tell. This is the dangerous category and it is the reason idempotency keys exist. It is also the only one of the three that cannot be classified from the response, because there is no response to classify. It has to be handled by design.
Figure 1 sets out all three. Most integrations classify the first two by accident, through whatever status code the target system happened to return, and never name the third at all. The third is the one that creates duplicates, and it is invisible in testing because tests rarely lose a response.

Figure 1: Three kinds of failure, sorted by whether retrying could work.
Gartner projects that by 2027, 70% of procurement intake requests will be AI-assisted.
Gartner was forecasting intake automation broadly rather than error handling. What follows is that failure classification stops being an edge case as volume shifts to automated paths, because a human noticing something looked wrong is no longer in the loop by default.
Download An Analyst Report – The Hackett Group Agentic AI in Procurement Adoption Index — 2026
How should each one be handled?
Permanent failures stop immediately and route to a person. The request is not broken infrastructure, it is a request that needs correcting, and the correction is usually trivial once someone sees it. The expensive part is the interval before anyone does. Speed of surfacing matters more than anything else here, because the fix is cheap and the waiting is not. A request held for a day over a cost center is a day nobody bought anything with.
Ambiguous failures retry with the same idempotency key, which is what makes the retry safe. Without the key, retrying an ambiguous failure is how duplicate purchase orders reach suppliers.
The classification has to be explicit in the integration contract. This classification belongs in the integration contract rather than in code. Inferring it from status codes works until a system returns a server error for what is actually a validation problem, which happens often enough to be planned for rather than discovered.
APQC measured transaction processing cost rather than exception handling. The bearing on this piece is that a duplicate or stranded request carries that cost twice while delivering nothing, which is the arithmetic that makes error design worth specifying rather than discovering.
Don’t let failed intake requests turn into manual fire drills. Streamline purchase order generation and eliminate integration errors with intelligent ERP synchronization. – Explore Zycus Purchase Order Management Software
Where do failed requests go?
To a dead letter queue with a named owner, and the second half of that sentence is the one that gets omitted. A queue is infrastructure. An owner is a commitment.
A dead letter queue is a holding place for messages that failed permanently and will not succeed on another attempt. Its purpose is to make failure visible rather than to store it, and the difference between those two is entirely a matter of what happens next. A queue nobody monitors is where requests go to be forgotten more thoroughly than if the integration had simply crashed, because a crash gets noticed and a silently growing queue does not.
Three things make it work. Someone owns it by name rather than by team. Its depth is monitored and alerts above a threshold. And there is a defined path back, so a corrected request can re-enter the flow rather than being re-keyed manually into the target system.
That third point is where most implementations stop. The queue exists, the alert exists, and the recovery is a person retyping the request into the ERP, which loses the audit trail connecting it to the original and makes the failure invisible to any later analysis of how often this happens.
Download A Free Whitepaper – Unlocking Deep Value in Intake Management with Merlin Intake
What does the requester see?
Something, as Figure 2 shows, and this is the requirement most often missed entirely.
A request that fails silently is worse than one that fails loudly, because the requester believes it is progressing. They wait, then chase, then bypass the process next time, and the bypass is permanent in a way the failure was not. One silent failure costs more adoption than several visible ones, and adoption is the thing intake was bought to improve.
The rule worth holding is that no failure state is invisible to the requester. They do not need the technical detail. They need to know that something needs attention, who has it, and roughly when to expect an answer. A request stuck in a dead letter queue with no requester-visible state is the single most damaging failure mode in intake, and it is entirely preventable at design time for almost no cost.

Figure 2: What silence costs, against what a visible state costs.
How do you test this before go-live?
Force each failure type deliberately, because none of the three appears in normal testing. Test environments have clean data and stable connections, which is precisely why they do not exercise error handling.
Submit a request with an invalid cost center and confirm it stops rather than retrying, reaches a named person, and shows the requester a state rather than nothing. Disconnect the ERP mid-transaction and confirm the retry uses the same key and produces one record rather than two. Fill the dead letter queue past its alert threshold and confirm somebody is actually notified.
Then take a request out of the dead letter queue, correct it, and confirm it re-enters the flow with its history intact rather than arriving as a new request with no connection to the original.
Merlin Intake integrates with ERP systems as part of the platform, which means these paths are defined rather than assembled per project. Whichever platform you run, run the four tests above before go-live rather than discovering the answers in production, where each one costs a real request.
Hackett studied the procurement agenda rather than integration failure. It matters here because manual recovery of failed requests is exactly the work a team with declining headcount cannot absorb, which is what makes the dead letter path a capacity question.
Related Reads
- Do you need structured intake or full eProcurement?
- How to Evaluate Procurement Intake Software: 7 Criteria That Matter
- What Is an Idempotency Key, and Why Does Your Intake-to-ERP Integration Need One?
- What must be true about your master data before intake goes live?
- What Should a Procurement Intake RFP Ask For?
Frequently asked questions
Q1. What is an error-handling strategy for system integration?
It is the explicit definition of how failures are classified, what happens to each class, where permanently failed messages go, who owns them, and what the originating user sees. Integrations without one work until conditions change, then fail in ways that are difficult to diagnose and easy to miss.
Q2. What is the difference between a transient and a permanent failure?
A transient failure is caused by temporary conditions, such as a timeout or a service restart, and retrying will usually succeed. A permanent failure is caused by the content of the request itself, such as an invalid cost center, and will fail identically on every attempt. Retrying permanent failures generates noise that buries genuine errors.
Q3. What is a dead letter queue?
A dead letter queue holds messages that failed permanently and cannot succeed on retry. Its purpose is to make failure visible and recoverable rather than to store it. Without a named owner, monitoring, and a defined path back into the flow, it becomes a place where requests are forgotten rather than handled.
Q4. How many times should a failed request retry?
Transient failures should retry a bounded number of times with increasing intervals between attempts, typically a handful rather than dozens. Permanent failures should not retry at all. Unbounded retry is not resilience, it applies sustained load to a system that is already struggling and delays the point at which a human sees the problem.
Q5. What should a requester see when an integration fails?
That something needs attention, who is handling it, and roughly when to expect an answer. They do not need technical detail. Silence is the damaging option, because a requester who believes their request is progressing waits, then chases, then bypasses the process on the next purchase.
Q6. How do you test integration error handling?
Force each failure type deliberately. Submit an invalid cost center and confirm it stops rather than retrying. Disconnect the target system mid-transaction and confirm the retry produces one record rather than two. Fill the dead letter queue past its threshold and confirm someone is notified. Then recover a message and confirm its history survives.
Q7. Who owns integration failures, procurement or IT?
IT owns the mechanism and procurement owns the consequences, since failed requests surface as procurement problems long before anyone traces them to an integration. In practice the dead letter queue needs a named owner in procurement operations, because deciding what a corrected request should say is a procurement judgment rather than a technical one.
Q8. Can failed requests be recovered automatically?
Transient failures recover automatically by design. Permanent failures cannot, because the request content needs correcting and only a person can decide what it should say. What can and should be automated is the path back: once corrected, a request should re-enter the flow with its original history intact rather than being re-keyed as a new request.





















































