First, let me explain why I chose this use case
Well, if you work in this field, you know refunds are one of the most repetitive and, at the same time, most delicate customer service processes. Every case demands the same investigation: finding the order, verifying who paid for it, checking whether the activity already took place and interpreting the policy. And every case touches money, so a mistake means a loss or a serious problem with your customer.
So what I did was build an AI Agent that does that work from start to finish inside Zendesk, with one basic condition: it acts on its own only when the decision is safe and provable. In every other case it asks for information, rejects with an explanation or hands the case to a human agent.
It has five pieces, each with a single responsibility. I didn’t add a user interface, since the team keeps working in my Zendesk instance, where two triggers notify my autonomous agent when a ticket is created with a refund form that I built, and when the customer replies to a pending case.
- Entry point
- Receives the event, verifies it is authentic and turns it into a strictly shaped object. An empty, mistyped or unexpected field is rejected and doesn’t contaminate the decision.
- Decision engine
- Two separate, versioned layers: the policy decides whether the case qualifies; the authority decides whether the system can act on its own. They are never mixed.
- Persistent flow
- Each case lives in a state machine saved in a database. If the customer replies days later, the case resumes where it left off.
- Commercial simulator
- My simulator has fictional orders, payments, activities and refunds that reproduce eleven scenarios, including duplicate charge, open dispute and mismatched identity.
- Deployment
- I used GitHub to store my autonomous agent’s code, Railway as the server to host it and SQLite as the database.
Design goals and principles
I defined six principles before writing any code:
- Automate only the safe cases.
- Fail closed. When in doubt, with missing data or an uncertain result, the system doesn’t execute the financial action.
- Keep people in control. Escalation is an intended outcome, with its own route, queue and message.
- Preserve evidence. Every state transition is recorded with its reason and the policy version applied.
- Avoid duplicate refunds. Refund creation is idempotent and an unknown result is reconciled before any retry.
- Communicate clearly in Zendesk. What the customer sees and what the team sees are two different things, but aligned with the use case.
Visually, this autonomous system works like this
Scroll horizontally to see the whole diagram.
gpt-4o-mini proposes one and the acceptance gate decides whether it is accepted. Only that number enters the deterministic engine, which is the only component that verifies facts and moves money.How the decision is made
The agent always goes through the same questions, in the same order, and stops at the first one it can’t answer with certainty.
- Is the order number included? If not, the case is put on hold and the customer is asked for it.
- Does that order exist? If it doesn’t show up, there is nothing to refund and the reason is explained.
- Is the sender the person who paid? If the requester’s email doesn’t match the order’s owner, the case goes to human review. It is neither rejected nor approved: it is handed over.
- Does it qualify under the policy? Four rules can make it eligible: withdrawal within seven days, cancellation of the activity by the academy, a confirmed duplicate charge and a confirmed technical failure. If several apply, there is a fixed order of precedence.
- Can the system execute on its own? A different question. A case can be eligible and still require a person: because of the amount, an open dispute or a manual-review flag.
- If not, reject with an explanation or escalate. Never stay silent.
Of course this flow had to have safeguards. Let me explain them visually and simply, take a look.
Scroll horizontally to see the whole diagram.
Safeguards and trust
Authentication and privacy
Every call carries a shared secret header, compared in a way that resists timing attacks. The flow tables store reason codes and fingerprints of the data, not the event body or the ticket text.
Identity and human review
If the sender doesn’t match the order’s owner, the case is assigned to a review group in Zendesk with an internal note and a structured reason. It is never approved or rejected on its own, and it never ends up without an owner.
Deterministic policy and authority
Rules in versioned code with tests. Above the configured monetary limit, the case requires a person even if it is perfectly eligible.
Duplicates and idempotency
Every event is recorded with a stable identifier, and refund creation uses a deterministic key. A resend doesn’t act again, and an unknown result is reconciled before any retry is allowed.
Evidence and atomic rollback
Every transition stores its reason, the policy version and a cryptographic fingerprint of the evidence. If a late step fails, the whole operation is rolled back: there are no half-written states.
Safe communication
Public texts are fixed templates. They never expose internal data or error traces, and they never promise an outcome that didn’t happen.
Now, what makes this system agentic?
Today the word “agentic” or “automation” is used for almost anything, so I’ll stop here to explain it properly because I don’t want to sell you smoke.
This system is agentic in an operational sense, meaning it closes the loop between perceiving and acting on its own.
- It perceives an event from my Zendesk instance without anyone invoking it manually.
- It keeps persistent case state across interactions separated by days.
- It gathers on its own the context it needs.
- It evaluates policy and authority, and chooses one action among several possible ones.
- It acts through Zendesk, adds comments and tags that I defined, and changes the ticket status.
- It executes or refuses a refund.
- It asks the customer for the missing information and waits for the answer.
- It escalates uncertain or unsafe cases to one of my human agents in Zendesk.
- It leaves evidence and an audit history of every transition.
It’s worth clarifying that no model takes part in the decision on eligibility, authority, identity or execution. Those decisions are deterministic and auditable, and the system doesn’t reason like a person to decide.
How does the conversational understanding layer work?
To understand my customer’s request, I implemented an understanding layer with gpt-4o-mini, which receives the customer’s last message and returns, through Structured Outputs, four fields: intent, stated reason, order number and confidence. The system validates the schema, the intent, the confidence and the format of the order number, and only the accepted order number reaches the deterministic engine. I kept the intent, the reason and the confidence only to decide whether to trust that text string, and then they are discarded.
Conditions for accepting an order number
- Only when there is nothing better. If the ticket already includes the order number, the model isn’t called at all, which saves me OpenAI costs.
- Refund-related intent and a minimum confidence of 0.80.
- Valid format. A malformed identifier is discarded, never corrected: guessing it would be guessing which payment to refund.
How do I measure confidence?
Confidence expresses how sure the model is that it extracted the text correctly, not whether what the customer says is true or whether the refund applies. They are two different questions, and only the first one is the model’s job.
What happens when it fails?
Things like a timeout, a network or API error, a response that doesn’t meet the schema, low confidence or an invalid credential all produce the same result: no order number, and the system keeps asking the customer for the exact identifier.
Here are some examples of how it works
I’m not including savings, satisfaction or volume metrics because I didn’t measure them: this is an MVP in demo form to show you that it’s possible.
| Scenario | Outcome | What was observed |
|---|---|---|
| Activity cancelled by the academy, eligible case | Simulated refund | End-to-end flow: refund recorded exactly once, ticket resolved and a public confirmation comment. |
| Order number missing | Information requested | The agent asked for the number and left the case on hold. When the customer replied, it resumed on the same case file instead of opening a new one. |
| The requester’s identity doesn’t match the order’s owner | Human review | No financial action. Case open and assigned to the review group, with an internal note and a structured reason. No misleading message to the customer. |
| Duplicate event | No repeated action | The second event was recognized as a duplicate: no second refund and no rewriting of the state. |
| Natural-language request, no order number in the form | Human review | The model extracted the intent, the stated reason and the order number with high confidence. The deterministic engine loaded the commercial facts, detected that the sender wasn’t the order’s owner and handed the case over. No refund authorized or executed, and no trigger loop. |
| Deployment restart with already-modified data | Service recovered | After fixing the startup, the service started with the volume’s data, reported the divergence and stayed healthy. No restart loops. |
What do the customer and my team see in Zendesk?
Depending on the outcome, a different action is taken.
| Outcome | Comment | Status | Effect |
|---|---|---|---|
| Refund approved | Public | Solved | Confirms the refund to the original payment method and clarifies that the crediting time depends on the bank. |
| Missing information | Public | Pending | Asks for the order number with a concrete example and leaves the case waiting for a reply. |
| Human review | Internal | Open | Private note with the reason and assignment to the review group. The customer doesn’t receive an automatic message that prejudges the case. |
| Not eligible | Public | Solved | Explains that the request doesn’t qualify under the policy, without closing the door to a review by a person. |
Current limitations
- Simulated payments. The executor isn’t connected to any real gateway; no money moves.
- Fictional data. Silbato, its academies, customers, orders and payments are made up.
- Understanding validated in controlled scenarios, not with real volume: it works and is active, but I haven’t measured its accuracy with customer traffic.
- No model decides about money. There is no generative component in eligibility, authority, identity or execution.
- SQLite and a single replica. MVP simplicity in exchange for not scaling horizontally.
- No reconciliation or appeals. There is no automatic accounting reconciliation or formal appeals flow.
- MVP validation, not certification. Nothing here has passed a security, compliance or risk review.
Now, how could this autonomous agent benefit you?
If you decide to take this use case to production, you could reduce handling time per case and avoid inconsistent answers, errors, duplicates, fraud and losses from incorrect refunds. It would also operate 24/7 and leave an audit trail of everything it does so your human team knows what is happening.
Does it save money?
Yes, and there’s no need to go around in circles: it saves money because it lowers the cost per refund and increases the team’s capacity without you having to hire more people.
Here’s an example of what the scenario looks like
Say you currently have 6 human agents dedicated to refunds. If you implement this autonomous agent, could you go from 6 agents to 1? The short answer is yes, but wait: you probably shouldn’t eliminate five positions right away. First you free up capacity for complex cases, retention and higher-value support. After the pilot you can decide whether to reduce by 5, or 4, or 3, etc., or redistribute them to more important areas that aren’t automated yet.
Here’s a simple formula to calculate the savings so it doesn’t stay just ideas
savings = automated cases × minutes saved × cost per minute
Measure during the pilot: volume, percentage resolved automatically, average time, escalations and errors. With that you’ll know how many FTEs it really frees up.
To conclude
The LLM I used understands human language, as is common by now; however, deterministic systems verify the facts and control financial actions. Separating those two jobs is what allows you to use a language model in a process that touches money without passing the risk on to it.
It’s also worth mentioning that four of the six validated scenarios end without moving money. That isn’t a limitation of this agent; it’s actually the purpose. A system that knows how to hand over a hard case well is more adoptable than one that tries to solve everything, in my opinion.
To wrap up, I just want to tell you that the hard part of an autonomous agent isn’t the automation. It’s the joint design of four things: the action the system can execute, the safeguards that delimit how far that capability goes, the evidence that lets you reconstruct every decision, and human escalation as a first-class outcome.
When all four are in place and you make it work, the conversation with your risk and finance team changes: you no longer have to ask yourself “do we trust the AI?”, but “do we agree with these limits we’ve given it?”.
This is a demo I built on a financial use case, one that involves money, like refunds, but don’t limit yourself: an autonomous or AI agent can cover many other scenarios or use cases.