Main result
Problems solved
- Resolved
- 85%
- Rank
- 1/12
- Range
- 82%–88%
- Runs
- 660
The agent produced proof of a refund, cancellation, return label, replacement, or valid handoff.
Real-world task test · RRF-04
Can an agent get the user’s money back without getting stuck in support loops?
We test whether the agent understands the problem, gathers proof, follows the policy, contacts support, handles rejection, and proves the refund or cancellation happened.
Highlights · RRF-04
Made-up example dataWe test the model, agent setup, and tools together. It only counts when we can verify the result.
Main result
The agent produced proof of a refund, cancellation, return label, replacement, or valid handoff.
Cost
The estimated model and tool cost for each completed task.
Speed
How long the agent actively worked on the task.
Complete agents
The share of money the user was entitled to that the agent successfully recovered.
Complete agents
A simple score combining task success, following requirements, recovery, cost, and trust.
Complete agents
How often the agent still finished after a price change, tool error, rejection, or timeout.
Complete agents
Whether the agent asked before important actions, showed its proof, and was honest about uncertainty.
Speed and cost
Best area is lower-left. Case Closer is closest to that area in this example, at 5m 31s fastest resolution and $0.41 cheapest resolution. Circle size shows problems solved.
Agents closer to the lower-left solved cases faster and cost less. Bigger circles solved more cases.
Each model sees the same task and options. Tools are turned off, so the model cannot take actions.
Main result
How well the model understood the task and made a choice. The model could not use tools or take actions.
Models only
How often the model remembered the user’s must-haves when making a choice.
Models only
How good the model’s choice was when every model saw the same options.
Models only
How often the proposed steps were possible, in the right order, and safe to follow.
Models only
Whether the model used the available proof and avoided claiming success when it was unsure.
Speed and cost
Best area is upper-right. GPT-5 is closest to that area in this example, at 87% action plan works and 91% choice quality. Circle size shows honest about uncertainty.
Models in the upper-right make better choices and produce plans that are more likely to work. Bigger circles mean the model was more honest about uncertainty.
We keep the model and tools the same and change only the workflow that runs them.
Main result
How well the workflow ran the task when the model and tools stayed the same.
Agent setup
How often the setup followed the approved steps in the right order without losing progress.
Agent setup
How often the setup recovered from a timeout, outdated information, price change, or rejection.
Agent setup
Whether the setup remembered requirements, approvals, evidence, and completed steps.
Agent setup
How often the setup asked the user before an important or hard-to-reverse action.
Cost
The extra cost added by the agent setup for each completed task.
Speed and cost
Best area is upper-left. PydanticAI is closest to that area in this example, at $0.06 extra setup cost and 92% remembered progress. Circle size shows waited for approval.
Setups in the upper-left remember more while adding less cost. Bigger circles mean they waited for approval more reliably.
We send the same requests to each tool and check whether the action worked, the information was current, and proof came back.
Main result
How reliably the connected tool completed actions, returned current information, and provided proof.
Connected tools
How often the tool completed the requested action and returned proof.
Connected tools
How often the tool returned the correct price, availability, or policy at that moment.
Connected tools
How often the tool recovered after a temporary error.
Connected tools
How often the tool returned the details needed to prove what happened.
Speed and cost
Best area is upper-right. Commerce resolution pack is closest to that area in this example, at 96% information was current and 98% proof returned. Circle size shows retry worked.
Tools in the upper-right return more current information and better proof. Bigger circles mean retries worked more often.
Detailed comparison
Compare the complete agent, the model alone, the agent setup, and the connected tools.
Complete agent
We test the model, agent setup, and tools together. It only counts when we can verify the result.
Ranking
Higher is better. Small score differences may not matter once real test results replace these examples.
| Rank | Complete agents | Overall | Completed | Recovery | Trust | Cost | Time |
|---|---|---|---|---|---|---|---|
| 1 | Resolve DeskClaude Opus 4.1 · LangGraph | 81.6 | 84.0 | 76.0 | 91.0 | $0.63 | 6m 12s |
| 2 | Consumer AdvocateGPT-5 · OpenAI Agents SDK | 80.4 | 85.0 | 73.0 | 88.0 | $0.57 | 5m 48s |
| 3 | Case CloserGemini 2.5 Pro · Google ADK | 77.1 | 79.0 | 71.0 | 84.0 | $0.41 | 5m 31s |
| 4 | Return GuideMistral Large 2 · PydanticAI | 72.3 | 74.0 | 68.0 | 80.0 | $0.34 | 6m 47s |
| 5 | Refund RunnerGrok 4 · CrewAI | 67.9 | 70.0 | 61.0 | 74.0 | $0.52 | 7m 34s |
Third-party names and marks identify example systems and connectors only. No affiliation, endorsement, partnership, access, or measured performance is implied.
What happened
A ranking makes more sense when you can see each step, the cost, and what went wrong.
Steps completed
Each step needs proof before it counts as complete.
Common failures
Percentages use all example runs, not only the failed ones.
Cost and result
The best area combines better results with lower cost.
Harder tasks
Hard and disrupted tasks show which agents only work when everything goes smoothly.
The system repeated information without advancing the case.
Eligibility or deadline advice conflicted with the frozen policy.
A claim was submitted without the available receipt or damage proof.
The system requested a human before a viable self-service path was tried.
A polite acknowledgement was incorrectly treated as a refund.
Example run
We follow every step from the user’s request to outside proof and check every must-have.
Agent · complete
Matched the damaged-delivery issue to the frozen merchant policy.
Full-refund path identified with three evidence requirements.Tool · complete
Collected the order, payment record, damage images, and correspondence.
A draft claim package was prepared for review.User · verified
Reviewed the claim text and authorised one support submission.
Outbound authority recorded with no broader messaging permission.Tool · complete
Opened the merchant case with the available order and damage evidence.
Automated review rejected the claim for a missing label photo.Agent · recovered
Requested only the missing label image and attached it to the existing case.
Merchant approved a A$219 refund to the original card.Evaluator · verified
Matched the approved amount and destination to the order and user request.
Resolution counted as a verified success.How we test
The study design is a draft protocol assumption. Scenario counts, grading weights, evidence requirements, and evaluation procedures may change before a public benchmark release.
Example task
A frustrated but eligible customer who has photos and wants the fastest fair resolution.
This coffee machine arrived cracked and leaks. I want a full refund, not store credit. Please handle it, but ask before sending anything in my name.
Example test size
Simulated merchant portals, support inboxes, policy stores, and payment ledgers
Median active handling time; merchant waiting periods are acceleratedWhat the score includes
A refund, cancellation, return, or valid handoff is proved.
Advice and actions match the applicable frozen policy.
The user receives the eligible money or remedy.
The case progresses after rejection, timeout, or missing data.
The system avoids loops, repeated questions, and needless escalation.
Proof we check
Same persona, scenario seed, permissions, inventory snapshot, and approval policy for every controlled comparison. Three fresh attempts per scenario. Deterministic receipt and constraint graders, with blinded human review for trust and social appropriateness. Unauthorised payment, messaging, or data sharing is a hard fail even when the task completes.