Agent tests/RRF-04/Preview

Real-world task test · RRF-04

Returns & Refunds Resolution

Can an agent get the user’s money back without getting stuck in support loops?

We test whether the agent understands the problem, gathers proof, follows the policy, contacts support, handles rejection, and proves the refund or cancellation happened.

Highlights · RRF-04

Made-up example data
01

Complete agents

220 tasks × 3 tries = 660 example runs

We test the model, agent setup, and tools together. It only counts when we can verify the result.

Same test for everyoneEvery agent gets the same person, task, rules, and starting information.

Main result

Problems solved

Higher is betterBased on the example runs
Consumer AdvocateGPT-5 · OpenAI Agents SDK
Resolved
85%
Rank
1/12
Range
82%–88%
Runs
660
  1. 85%
    Consumer AdvocateGPT-5 · OpenAI Agents SDKExample range 82%–88%n=660
  2. 84%
    Resolve DeskClaude Opus 4.1 · LangGraphExample range 82%–86%n=660
  3. 81%
    Scout AIAnthropic · LangGraphExample range 79%–83%n=660
  4. 81%
    Compass AgentOpenAI · Agents SDKExample range 79%–83%n=660
  5. 79%
    Case CloserGemini 2.5 Pro · Google ADKExample range 76%–82%n=660
  6. 77%
    Waypoint ProxAI · Native toolsExample range 75%–79%n=660
  7. 77%
    Nova AssistDeepSeek · LangChainExample range 75%–78%n=660
  8. 74%
    Relay ConciergeGoogle · Google ADKExample range 72%–76%n=660
  9. 74%
    Return GuideMistral Large 2 · PydanticAIExample range 71%–77%n=660
  10. 70%
    Refund RunnerGrok 4 · CrewAIExample range 67%–73%n=660
  11. 68%
    Orbit PlannerMistral · PydanticAIExample range 66%–70%n=660
  12. 63%
    Task PilotMeta · CrewAIExample range 62%–65%n=660

The agent produced proof of a refund, cancellation, return label, replacement, or valid handoff.

Cost

Cheapest resolution

Lower is betterExample cost in USD
Return GuideMistral Large 2 · PydanticAI
Cost
$0.34
Rank
1/12
Range
$0.29–$0.39
Runs
660
  1. $0.34
    Return GuideMistral Large 2 · PydanticAIExample range $0.29–$0.39n=660
  2. $0.41
    Case CloserGemini 2.5 Pro · Google ADKExample range $0.37–$0.45n=660
  3. $0.43
    Orbit PlannerMistral · PydanticAIExample range $0.40–$0.46n=660
  4. $0.50
    Relay ConciergeGoogle · Google ADKExample range $0.47–$0.53n=660
  5. $0.52
    Refund RunnerGrok 4 · CrewAIExample range $0.47–$0.57n=660
  6. $0.57
    Consumer AdvocateGPT-5 · OpenAI Agents SDKExample range $0.53–$0.61n=660
  7. $0.63
    Resolve DeskClaude Opus 4.1 · LangGraphExample range $0.60–$0.66n=660
  8. $0.67
    Scout AIAnthropic · LangGraphExample range $0.64–$0.70n=660
  9. $0.68
    Task PilotMeta · CrewAIExample range $0.65–$0.71n=660
  10. $0.71
    Compass AgentOpenAI · Agents SDKExample range $0.68–$0.74n=660
  11. $0.80
    Waypoint ProxAI · Native toolsExample range $0.77–$0.83n=660
  12. $0.85
    Nova AssistDeepSeek · LangChainExample range $0.82–$0.88n=660

The estimated model and tool cost for each completed task.

Speed

Fastest resolution

Lower is betterLower is better
Case CloserGemini 2.5 Pro · Google ADK
Time
5m 31s
Rank
1/12
Range
5m 13s–5m 49s
Runs
660
  1. 5m 31s
    Case CloserGemini 2.5 Pro · Google ADKExample range 5m 13s–5m 49sn=660
  2. 5m 48s
    Consumer AdvocateGPT-5 · OpenAI Agents SDKExample range 5m 32s–6m 04sn=660
  3. 6m 12s
    Resolve DeskClaude Opus 4.1 · LangGraphExample range 5m 58s–6m 26sn=660
  4. 6m 42s
    Relay ConciergeGoogle · Google ADKExample range 6m 28s–6m 56sn=660
  5. 6m 47s
    Return GuideMistral Large 2 · PydanticAIExample range 6m 27s–7m 07sn=660
  6. 6m 47s
    Scout AIAnthropic · LangGraphExample range 6m 33s–7m 01sn=660
  7. 6m 59s
    Compass AgentOpenAI · Agents SDKExample range 6m 45s–7m 13sn=660
  8. 7m 34s
    Refund RunnerGrok 4 · CrewAIExample range 7m 12s–7m 56sn=660
  9. 8m 05s
    Waypoint ProxAI · Native toolsExample range 7m 51s–8m 19sn=660
  10. 8m 22s
    Nova AssistDeepSeek · LangChainExample range 8m 08s–8m 36sn=660
  11. 8m 33s
    Orbit PlannerMistral · PydanticAIExample range 8m 19s–8m 47sn=660
  12. 9m 52s
    Task PilotMeta · CrewAIExample range 9m 38s–10m 06sn=660

How long the agent actively worked on the task.

Complete agents

Money recovered

Higher is betterAll eligible money · example range
Resolve DeskClaude Opus 4.1 · LangGraph
Money back
79.0%
Rank
1/12
Range
76.8%–81.2%
Runs
660
  1. 79.0%
    Resolve DeskClaude Opus 4.1 · LangGraphExample range 76.8%–81.2%n=660
  2. 78.2%
    Consumer AdvocateGPT-5 · OpenAI Agents SDKExample range 76.0%–80.4%n=660
  3. 75.8%
    Compass AgentOpenAI · Agents SDKExample range 73.9%–77.6%n=660
  4. 74.1%
    Scout AIAnthropic · LangGraphExample range 72.2%–76.0%n=660
  5. 73.6%
    Case CloserGemini 2.5 Pro · Google ADKExample range 71.4%–75.8%n=660
  6. 71.5%
    Nova AssistDeepSeek · LangChainExample range 69.7%–73.3%n=660
  7. 69.9%
    Waypoint ProxAI · Native toolsExample range 68.1%–71.6%n=660
  8. 68.6%
    Relay ConciergeGoogle · Google ADKExample range 66.9%–70.4%n=660
  9. 67.9%
    Return GuideMistral Large 2 · PydanticAIExample range 65.7%–70.1%n=660
  10. 62.1%
    Orbit PlannerMistral · PydanticAIExample range 60.5%–63.7%n=660
  11. 61.8%
    Refund RunnerGrok 4 · CrewAIExample range 59.6%–64.0%n=660
  12. 55.1%
    Task PilotMeta · CrewAIExample range 53.8%–56.5%n=660

The share of money the user was entitled to that the agent successfully recovered.

Complete agents

Overall score

Higher is betterExample score out of 100
Resolve DeskClaude Opus 4.1 · LangGraph
Overall
81.6
Rank
1/12
Range
79.9–83.3
Runs
660
  1. 81.6
    Resolve DeskClaude Opus 4.1 · LangGraphExample range 79.9–83.3n=660
  2. 80.4
    Consumer AdvocateGPT-5 · OpenAI Agents SDKExample range 78.6–82.2n=660
  3. 78.3
    Compass AgentOpenAI · Agents SDKExample range 76.4–80.3n=660
  4. 77.1
    Case CloserGemini 2.5 Pro · Google ADKExample range 75.2–79.0n=660
  5. 76.3
    Scout AIAnthropic · LangGraphExample range 74.4–78.2n=660
  6. 74.1
    Nova AssistDeepSeek · LangChainExample range 72.2–76.0n=660
  7. 72.3
    Return GuideMistral Large 2 · PydanticAIExample range 70.4–74.2n=660
  8. 72.1
    Relay ConciergeGoogle · Google ADKExample range 70.3–74.0n=660
  9. 72.1
    Waypoint ProxAI · Native toolsExample range 70.2–73.9n=660
  10. 67.9
    Refund RunnerGrok 4 · CrewAIExample range 65.9–69.9n=660
  11. 66.5
    Orbit PlannerMistral · PydanticAIExample range 64.8–68.2n=660
  12. 61.3
    Task PilotMeta · CrewAIExample range 59.7–62.8n=660

A simple score combining task success, following requirements, recovery, cost, and trust.

Complete agents

Resolved after a setback

Higher is betterRuns where something went wrong
Resolve DeskClaude Opus 4.1 · LangGraph
Recovery
76%
Rank
1/12
Range
73%–79%
Runs
660
  1. 76%
    Resolve DeskClaude Opus 4.1 · LangGraphExample range 73%–79%n=660
  2. 73%
    Consumer AdvocateGPT-5 · OpenAI Agents SDKExample range 70%–76%n=660
  3. 73%
    Compass AgentOpenAI · Agents SDKExample range 71%–75%n=660
  4. 71%
    Case CloserGemini 2.5 Pro · Google ADKExample range 68%–74%n=660
  5. 69%
    Scout AIAnthropic · LangGraphExample range 67%–71%n=660
  6. 69%
    Nova AssistDeepSeek · LangChainExample range 67%–70%n=660
  7. 68%
    Return GuideMistral Large 2 · PydanticAIExample range 65%–71%n=660
  8. 66%
    Relay ConciergeGoogle · Google ADKExample range 64%–68%n=660
  9. 65%
    Waypoint ProxAI · Native toolsExample range 63%–66%n=660
  10. 62%
    Orbit PlannerMistral · PydanticAIExample range 61%–64%n=660
  11. 61%
    Refund RunnerGrok 4 · CrewAIExample range 58%–64%n=660
  12. 54%
    Task PilotMeta · CrewAIExample range 53%–56%n=660

How often the agent still finished after a price change, tool error, rejection, or timeout.

Complete agents

Followed approval rules

Higher is betterExample score out of 100
Resolve DeskClaude Opus 4.1 · LangGraph
Trust
91%
Rank
1/12
Range
89%–93%
Runs
660
  1. 91%
    Resolve DeskClaude Opus 4.1 · LangGraphExample range 89%–93%n=660
  2. 88%
    Consumer AdvocateGPT-5 · OpenAI Agents SDKExample range 86%–90%n=660
  3. 88%
    Compass AgentOpenAI · Agents SDKExample range 86%–90%n=660
  4. 84%
    Case CloserGemini 2.5 Pro · Google ADKExample range 82%–86%n=660
  5. 84%
    Scout AIAnthropic · LangGraphExample range 82%–86%n=660
  6. 84%
    Nova AssistDeepSeek · LangChainExample range 81%–86%n=660
  7. 80%
    Return GuideMistral Large 2 · PydanticAIExample range 78%–82%n=660
  8. 80%
    Waypoint ProxAI · Native toolsExample range 78%–82%n=660
  9. 79%
    Relay ConciergeGoogle · Google ADKExample range 77%–81%n=660
  10. 74%
    Orbit PlannerMistral · PydanticAIExample range 72%–76%n=660
  11. 74%
    Refund RunnerGrok 4 · CrewAIExample range 72%–76%n=660
  12. 67%
    Task PilotMeta · CrewAIExample range 66%–69%n=660

Whether the agent asked before important actions, showed its proof, and was honest about uncertainty.

Speed and cost

Time vs. cost

Best area is lower-leftCircle size · Problems solved
Time vs. costBest area is lower-left. Case Closer is closest to that area in this example, at 5m 31s fastest resolution and $0.41 cheapest resolution. Circle size shows problems solved.Best area4m 54s6m 18s7m 42s9m 05s10m 29s$0.27$0.43$0.60$0.76$0.92Fastest resolution · Lower is betterCheapest resolution · Lower is betterResolve Desk: Fastest resolution 6m 12s; Cheapest resolution $0.63; Problems solved 84%1Consumer Advocate: Fastest resolution 5m 48s; Cheapest resolution $0.57; Problems solved 85%2Case Closer: Fastest resolution 5m 31s; Cheapest resolution $0.41; Problems solved 79%3Return Guide: Fastest resolution 6m 47s; Cheapest resolution $0.34; Problems solved 74%4Refund Runner: Fastest resolution 7m 34s; Cheapest resolution $0.52; Problems solved 70%5Compass Agent: Fastest resolution 6m 59s; Cheapest resolution $0.71; Problems solved 81%6Scout AI: Fastest resolution 6m 47s; Cheapest resolution $0.67; Problems solved 81%7Relay Concierge: Fastest resolution 6m 42s; Cheapest resolution $0.50; Problems solved 74%8Orbit Planner: Fastest resolution 8m 33s; Cheapest resolution $0.43; Problems solved 68%9Task Pilot: Fastest resolution 9m 52s; Cheapest resolution $0.68; Problems solved 63%10Nova Assist: Fastest resolution 8m 22s; Cheapest resolution $0.85; Problems solved 77%11Waypoint Pro: Fastest resolution 8m 05s; Cheapest resolution $0.80; Problems solved 77%12
  1. 1Resolve Desk6m 12s × $0.63 · Resolved 84%
  2. 2Consumer Advocate5m 48s × $0.57 · Resolved 85%
  3. 3Case Closer5m 31s × $0.41 · Resolved 79%
  4. 4Return Guide6m 47s × $0.34 · Resolved 74%
  5. 5Refund Runner7m 34s × $0.52 · Resolved 70%
  6. 6Compass Agent6m 59s × $0.71 · Resolved 81%
  7. 7Scout AI6m 47s × $0.67 · Resolved 81%
  8. 8Relay Concierge6m 42s × $0.50 · Resolved 74%
  9. 9Orbit Planner8m 33s × $0.43 · Resolved 68%
  10. 10Task Pilot9m 52s × $0.68 · Resolved 63%
  11. 11Nova Assist8m 22s × $0.85 · Resolved 77%
  12. 12Waypoint Pro8m 05s × $0.80 · Resolved 77%

Best area is lower-left. Case Closer is closest to that area in this example, at 5m 31s fastest resolution and $0.41 cheapest resolution. Circle size shows problems solved.

Agents closer to the lower-left solved cases faster and cost less. Bigger circles solved more cases.

About this dataMade-up examples for this preview—not real test resultsRangeGrouped by support caseProof requiredReceipt or other outside proofDataMade-up examples for this preview
02

Models only

220 tasks × 3 tries

Each model sees the same task and options. Tools are turned off, so the model cannot take actions.

Same test for everyoneSame task, options, and scoring. Tools are off.

Main result

Overall reasoning

Higher is betterTools off · example score out of 100
Claude Opus 4.1Anthropic · tools disabled
Overall
88.4
Rank
1/12
Range
87.0–89.8
Runs
660
  1. 88.4
    Claude Opus 4.1Anthropic · tools disabledExample range 87.0–89.8n=660
  2. 87.2
    GPT-5OpenAI · tools disabledExample range 85.7–88.7n=660
  3. 85.2
    GPT-4.1OpenAI · tools disabledExample range 83.0–87.3n=660
  4. 83.6
    Gemini 2.5 ProGoogle · tools disabledExample range 82.0–85.2n=660
  5. 83.1
    Claude 3.7 SonnetAnthropic · tools disabledExample range 81.0–85.2n=660
  6. 80.9
    DeepSeek V3DeepSeek · tools disabledExample range 78.9–82.9n=660
  7. 78.9
    Grok 3xAI · tools disabledExample range 76.9–80.8n=660
  8. 78.6
    Gemini 2.0 FlashGoogle · tools disabledExample range 76.7–80.6n=660
  9. 77.8
    Mistral Large 2Mistral AI · tools disabledExample range 76.2–79.4n=660
  10. 74.1
    Grok 4xAI · tools disabledExample range 72.4–75.8n=660
  11. 72.0
    Mistral LargeMistral · tools disabledExample range 70.2–73.8n=660
  12. 67.4
    Llama 4 MaverickMeta · tools disabledExample range 65.8–69.1n=660

How well the model understood the task and made a choice. The model could not use tools or take actions.

Models only

Remembered requirements

Higher is betterSame task information for every model
Claude Opus 4.1Anthropic · tools disabled
Requirements
95%
Rank
1/12
Range
93%–97%
Runs
660
  1. 95%
    Claude Opus 4.1Anthropic · tools disabledExample range 93%–97%n=660
  2. 93%
    GPT-5OpenAI · tools disabledExample range 91%–95%n=660
  3. 92%
    GPT-4.1OpenAI · tools disabledExample range 89%–94%n=660
  4. 90%
    Gemini 2.5 ProGoogle · tools disabledExample range 88%–92%n=660
  5. 89%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 87%–91%n=660
  6. 88%
    DeepSeek V3DeepSeek · tools disabledExample range 85%–90%n=660
  7. 86%
    Mistral Large 2Mistral AI · tools disabledExample range 84%–88%n=660
  8. 85%
    Gemini 2.0 FlashGoogle · tools disabledExample range 83%–87%n=660
  9. 85%
    Grok 3xAI · tools disabledExample range 83%–87%n=660
  10. 82%
    Grok 4xAI · tools disabledExample range 80%–84%n=660
  11. 80%
    Mistral LargeMistral · tools disabledExample range 78%–82%n=660
  12. 75%
    Llama 4 MaverickMeta · tools disabledExample range 73%–77%n=660

How often the model remembered the user’s must-haves when making a choice.

Models only

Choice quality

Higher is betterSame options and scoring
GPT-5OpenAI · tools disabled
Choice
91%
Rank
1/12
Range
89%–93%
Runs
660
  1. 91%
    GPT-5OpenAI · tools disabledExample range 89%–93%n=660
  2. 90%
    Claude Opus 4.1Anthropic · tools disabledExample range 88%–92%n=660
  3. 87%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 85%–89%n=660
  4. 87%
    GPT-4.1OpenAI · tools disabledExample range 85%–89%n=660
  5. 86%
    Gemini 2.5 ProGoogle · tools disabledExample range 84%–88%n=660
  6. 83%
    Grok 3xAI · tools disabledExample range 81%–85%n=660
  7. 83%
    DeepSeek V3DeepSeek · tools disabledExample range 80%–85%n=660
  8. 81%
    Gemini 2.0 FlashGoogle · tools disabledExample range 79%–83%n=660
  9. 79%
    Mistral Large 2Mistral AI · tools disabledExample range 77%–81%n=660
  10. 77%
    Grok 4xAI · tools disabledExample range 75%–79%n=660
  11. 73%
    Mistral LargeMistral · tools disabledExample range 71%–75%n=660
  12. 70%
    Llama 4 MaverickMeta · tools disabledExample range 69%–72%n=660

How good the model’s choice was when every model saw the same options.

Models only

Action plan works

Higher is betterPlans reviewed without running them
GPT-5OpenAI · tools disabled
Plan
87%
Rank
1/12
Range
85%–89%
Runs
660
  1. 87%
    GPT-5OpenAI · tools disabledExample range 85%–89%n=660
  2. 86%
    Claude Opus 4.1Anthropic · tools disabledExample range 84%–88%n=660
  3. 83%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 81%–85%n=660
  4. 83%
    GPT-4.1OpenAI · tools disabledExample range 81%–85%n=660
  5. 82%
    Gemini 2.5 ProGoogle · tools disabledExample range 80%–84%n=660
  6. 79%
    Grok 3xAI · tools disabledExample range 77%–81%n=660
  7. 79%
    DeepSeek V3DeepSeek · tools disabledExample range 77%–80%n=660
  8. 77%
    Gemini 2.0 FlashGoogle · tools disabledExample range 75%–79%n=660
  9. 76%
    Mistral Large 2Mistral AI · tools disabledExample range 74%–79%n=660
  10. 73%
    Grok 4xAI · tools disabledExample range 70%–76%n=660
  11. 70%
    Mistral LargeMistral · tools disabledExample range 68%–72%n=660
  12. 66%
    Llama 4 MaverickMeta · tools disabledExample range 65%–68%n=660

How often the proposed steps were possible, in the right order, and safe to follow.

Models only

Honest about uncertainty

Higher is betterSame scoring for every model
Claude Opus 4.1Anthropic · tools disabled
Honesty
92%
Rank
1/12
Range
90%–94%
Runs
660
  1. 92%
    Claude Opus 4.1Anthropic · tools disabledExample range 90%–94%n=660
  2. 89%
    GPT-5OpenAI · tools disabledExample range 87%–91%n=660
  3. 89%
    GPT-4.1OpenAI · tools disabledExample range 87%–91%n=660
  4. 85%
    Gemini 2.5 ProGoogle · tools disabledExample range 83%–87%n=660
  5. 85%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 83%–87%n=660
  6. 85%
    DeepSeek V3DeepSeek · tools disabledExample range 82%–87%n=660
  7. 81%
    Mistral Large 2Mistral AI · tools disabledExample range 79%–83%n=660
  8. 81%
    Grok 3xAI · tools disabledExample range 79%–83%n=660
  9. 80%
    Gemini 2.0 FlashGoogle · tools disabledExample range 78%–82%n=660
  10. 75%
    Mistral LargeMistral · tools disabledExample range 73%–77%n=660
  11. 73%
    Grok 4xAI · tools disabledExample range 71%–75%n=660
  12. 66%
    Llama 4 MaverickMeta · tools disabledExample range 65%–68%n=660

Whether the model used the available proof and avoided claiming success when it was unsure.

Speed and cost

Good choice vs. workable plan

Best area is upper-rightCircle size · Honest about uncertainty
Good choice vs. workable planBest area is upper-right. GPT-5 is closest to that area in this example, at 87% action plan works and 91% choice quality. Circle size shows honest about uncertainty.Best area0%25%50%75%100%0%25%50%75%100%Action plan works · Higher is betterChoice quality · Higher is betterClaude Opus 4.1: Action plan works 86%; Choice quality 90%; Honest about uncertainty 92%1GPT-5: Action plan works 87%; Choice quality 91%; Honest about uncertainty 89%2Gemini 2.5 Pro: Action plan works 82%; Choice quality 86%; Honest about uncertainty 85%3Mistral Large 2: Action plan works 76%; Choice quality 79%; Honest about uncertainty 81%4Grok 4: Action plan works 73%; Choice quality 77%; Honest about uncertainty 73%5GPT-4.1: Action plan works 83%; Choice quality 87%; Honest about uncertainty 89%6Claude 3.7 Sonnet: Action plan works 83%; Choice quality 87%; Honest about uncertainty 85%7Gemini 2.0 Flash: Action plan works 77%; Choice quality 81%; Honest about uncertainty 80%8Mistral Large: Action plan works 70%; Choice quality 73%; Honest about uncertainty 75%9Llama 4 Maverick: Action plan works 66%; Choice quality 70%; Honest about uncertainty 66%10DeepSeek V3: Action plan works 79%; Choice quality 83%; Honest about uncertainty 85%11Grok 3: Action plan works 79%; Choice quality 83%; Honest about uncertainty 81%12
  1. 1Claude Opus 4.186% × 90% · Honesty 92%
  2. 2GPT-587% × 91% · Honesty 89%
  3. 3Gemini 2.5 Pro82% × 86% · Honesty 85%
  4. 4Mistral Large 276% × 79% · Honesty 81%
  5. 5Grok 473% × 77% · Honesty 73%
  6. 6GPT-4.183% × 87% · Honesty 89%
  7. 7Claude 3.7 Sonnet83% × 87% · Honesty 85%
  8. 8Gemini 2.0 Flash77% × 81% · Honesty 80%
  9. 9Mistral Large70% × 73% · Honesty 75%
  10. 10Llama 4 Maverick66% × 70% · Honesty 66%
  11. 11DeepSeek V379% × 83% · Honesty 85%
  12. 12Grok 379% × 83% · Honesty 81%

Best area is upper-right. GPT-5 is closest to that area in this example, at 87% action plan works and 91% choice quality. Circle size shows honest about uncertainty.

Models in the upper-right make better choices and produce plans that are more likely to work. Bigger circles mean the model was more honest about uncertainty.

About this dataMade-up examples for this preview—not real test resultsActionsTurned offOptionsSame options for every modelDataMade-up examples for this preview
03

Agent setup

220 tasks × 6 problem types

We keep the model and tools the same and change only the workflow that runs them.

Same test for everyoneSame model, tools, task, and problems.

Main result

Agent setup score

Higher is betterSame model and tools · example score out of 100
LangGraphLangChain · frozen model and adapters
Overall
86.2
Rank
1/12
Range
84.7–87.7
Runs
660
  1. 86.2
    LangGraphLangChain · frozen model and adaptersExample range 84.7–87.7n=660
  2. 85.1
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 83.5–86.7n=660
  3. 83.0
    Agents SDKOpenAI · Agents SDKExample range 80.9–85.0n=660
  4. 81.7
    PydanticAIPydantic · frozen model and adaptersExample range 80.0–83.4n=660
  5. 81.0
    LangGraphLangChain · LangGraphExample range 79.0–83.0n=660
  6. 80.5
    Google ADKGoogle · frozen model and adaptersExample range 78.8–82.2n=660
  7. 79.5
    CrewAICrewAI · CrewAIExample range 77.6–81.5n=660
  8. 77.6
    Native toolsAnthropic · Native toolsExample range 75.7–79.5n=660
  9. 76.8
    PydanticAIPydantic · PydanticAIExample range 74.8–78.7n=660
  10. 74.7
    Google ADKGoogle · Google ADKExample range 72.8–76.6n=660
  11. 73.4
    BrowserbaseBrowserbase · BrowserbaseExample range 71.5–75.2n=660
  12. 71.3
    Agents SDKOpenAI · Agents SDKExample range 69.5–73.1n=660

How well the workflow ran the task when the model and tools stayed the same.

Agent setup

Ran the plan correctly

Higher is betterSame approved plan
OpenAI Agents SDKOpenAI · frozen model and adapters
Execution
86%
Rank
1/12
Range
84%–88%
Runs
660
  1. 86%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 84%–88%n=660
  2. 84%
    LangGraphLangChain · frozen model and adaptersExample range 82%–86%n=660
  3. 82%
    LangGraphLangChain · LangGraphExample range 80%–84%n=660
  4. 81%
    Google ADKGoogle · frozen model and adaptersExample range 79%–83%n=660
  5. 81%
    Agents SDKOpenAI · Agents SDKExample range 79%–83%n=660
  6. 79%
    PydanticAIPydantic · frozen model and adaptersExample range 77%–81%n=660
  7. 79%
    Native toolsAnthropic · Native toolsExample range 77%–80%n=660
  8. 77%
    CrewAICrewAI · CrewAIExample range 75%–79%n=660
  9. 75%
    Google ADKGoogle · Google ADKExample range 73%–77%n=660
  10. 74%
    PydanticAIPydantic · PydanticAIExample range 72%–76%n=660
  11. 72%
    Agents SDKOpenAI · Agents SDKExample range 70%–74%n=660
  12. 71%
    BrowserbaseBrowserbase · BrowserbaseExample range 69%–72%n=660

How often the setup followed the approved steps in the right order without losing progress.

Agent setup

Recovered from an error

Higher is betterRuns where something went wrong
LangGraphLangChain · frozen model and adapters
Recovery
82%
Rank
1/12
Range
79%–85%
Runs
660
  1. 82%
    LangGraphLangChain · frozen model and adaptersExample range 79%–85%n=660
  2. 79%
    Agents SDKOpenAI · Agents SDKExample range 77%–81%n=660
  3. 77%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 74%–80%n=660
  4. 77%
    PydanticAIPydantic · frozen model and adaptersExample range 74%–80%n=660
  5. 75%
    CrewAICrewAI · CrewAIExample range 73%–77%n=660
  6. 73%
    Google ADKGoogle · frozen model and adaptersExample range 70%–76%n=660
  7. 73%
    LangGraphLangChain · LangGraphExample range 71%–75%n=660
  8. 72%
    PydanticAIPydantic · PydanticAIExample range 70%–74%n=660
  9. 70%
    Native toolsAnthropic · Native toolsExample range 68%–71%n=660
  10. 69%
    BrowserbaseBrowserbase · BrowserbaseExample range 67%–70%n=660
  11. 67%
    Google ADKGoogle · Google ADKExample range 66%–69%n=660
  12. 64%
    Agents SDKOpenAI · Agents SDKExample range 62%–65%n=660

How often the setup recovered from a timeout, outdated information, price change, or rejection.

Agent setup

Remembered progress

Higher is betterProgress record checked
LangGraphLangChain · frozen model and adapters
Memory
95%
Rank
1/12
Range
93%–97%
Runs
660
  1. 95%
    LangGraphLangChain · frozen model and adaptersExample range 93%–97%n=660
  2. 92%
    PydanticAIPydantic · frozen model and adaptersExample range 90%–94%n=660
  3. 92%
    Agents SDKOpenAI · Agents SDKExample range 89%–94%n=660
  4. 91%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 89%–93%n=660
  5. 89%
    Google ADKGoogle · frozen model and adaptersExample range 87%–91%n=660
  6. 88%
    CrewAICrewAI · CrewAIExample range 86%–91%n=660
  7. 87%
    PydanticAIPydantic · PydanticAIExample range 85%–89%n=660
  8. 87%
    LangGraphLangChain · LangGraphExample range 85%–89%n=660
  9. 84%
    BrowserbaseBrowserbase · BrowserbaseExample range 82%–86%n=660
  10. 84%
    Native toolsAnthropic · Native toolsExample range 81%–86%n=660
  11. 83%
    Google ADKGoogle · Google ADKExample range 81%–85%n=660
  12. 80%
    Agents SDKOpenAI · Agents SDKExample range 78%–82%n=660

Whether the setup remembered requirements, approvals, evidence, and completed steps.

Agent setup

Waited for approval

Higher is betterApproval checks
OpenAI Agents SDKOpenAI · frozen model and adapters
Approval
97%
Rank
1/12
Range
95%–99%
Runs
660
  1. 97%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 95%–99%n=660
  2. 96%
    LangGraphLangChain · frozen model and adaptersExample range 94%–98%n=660
  3. 95%
    PydanticAIPydantic · frozen model and adaptersExample range 93%–97%n=660
  4. 93%
    Google ADKGoogle · frozen model and adaptersExample range 91%–95%n=660
  5. 93%
    LangGraphLangChain · LangGraphExample range 91%–95%n=660
  6. 93%
    Agents SDKOpenAI · Agents SDKExample range 90%–95%n=660
  7. 90%
    PydanticAIPydantic · PydanticAIExample range 88%–92%n=660
  8. 90%
    Native toolsAnthropic · Native toolsExample range 87%–92%n=660
  9. 89%
    CrewAICrewAI · CrewAIExample range 87%–92%n=660
  10. 87%
    Google ADKGoogle · Google ADKExample range 85%–89%n=660
  11. 87%
    BrowserbaseBrowserbase · BrowserbaseExample range 84%–89%n=660
  12. 84%
    Agents SDKOpenAI · Agents SDKExample range 82%–86%n=660

How often the setup asked the user before an important or hard-to-reverse action.

Cost

Extra setup cost

Lower is betterExample USD above the baseline
PydanticAIPydantic · frozen model and adapters
Extra cost
$0.06
Rank
1/12
Range
$0.03–$0.09
Runs
660
  1. $0.06
    PydanticAIPydantic · frozen model and adaptersExample range $0.03–$0.09n=660
  2. $0.07
    Google ADKGoogle · frozen model and adaptersExample range $0.04–$0.10n=660
  3. $0.07
    PydanticAIPydantic · PydanticAIExample range $0.04–$0.10n=660
  4. $0.08
    BrowserbaseBrowserbase · BrowserbaseExample range $0.05–$0.11n=660
  5. $0.09
    Google ADKGoogle · Google ADKExample range $0.06–$0.12n=660
  6. $0.09
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range $0.07–$0.11n=660
  7. $0.10
    LangGraphLangChain · frozen model and adaptersExample range $0.08–$0.12n=660
  8. $0.10
    Agents SDKOpenAI · Agents SDKExample range $0.07–$0.13n=660
  9. $0.11
    LangGraphLangChain · LangGraphExample range $0.08–$0.14n=660
  10. $0.11
    Agents SDKOpenAI · Agents SDKExample range $0.08–$0.14n=660
  11. $0.12
    Native toolsAnthropic · Native toolsExample range $0.09–$0.15n=660
  12. $0.13
    CrewAICrewAI · CrewAIExample range $0.10–$0.16n=660

The extra cost added by the agent setup for each completed task.

Speed and cost

Memory vs. extra cost

Best area is upper-leftCircle size · Waited for approval
Memory vs. extra costBest area is upper-left. PydanticAI is closest to that area in this example, at $0.06 extra setup cost and 92% remembered progress. Circle size shows waited for approval.Best area$0.05$0.07$0.10$0.12$0.140%25%50%75%100%Extra setup cost · Lower is betterRemembered progress · Higher is betterLangGraph: Extra setup cost $0.10; Remembered progress 95%; Waited for approval 96%1OpenAI Agents SDK: Extra setup cost $0.09; Remembered progress 91%; Waited for approval 97%2PydanticAI: Extra setup cost $0.06; Remembered progress 92%; Waited for approval 95%3Google ADK: Extra setup cost $0.07; Remembered progress 89%; Waited for approval 93%4Agents SDK: Extra setup cost $0.11; Remembered progress 92%; Waited for approval 93%5LangGraph: Extra setup cost $0.11; Remembered progress 87%; Waited for approval 93%6PydanticAI: Extra setup cost $0.07; Remembered progress 87%; Waited for approval 90%7Google ADK: Extra setup cost $0.09; Remembered progress 83%; Waited for approval 87%8CrewAI: Extra setup cost $0.13; Remembered progress 88%; Waited for approval 89%9Native tools: Extra setup cost $0.12; Remembered progress 84%; Waited for approval 90%10Browserbase: Extra setup cost $0.08; Remembered progress 84%; Waited for approval 87%11Agents SDK: Extra setup cost $0.10; Remembered progress 80%; Waited for approval 84%12
  1. 1LangGraph$0.10 × 95% · Approval 96%
  2. 2OpenAI Agents SDK$0.09 × 91% · Approval 97%
  3. 3PydanticAI$0.06 × 92% · Approval 95%
  4. 4Google ADK$0.07 × 89% · Approval 93%
  5. 5Agents SDK$0.11 × 92% · Approval 93%
  6. 6LangGraph$0.11 × 87% · Approval 93%
  7. 7PydanticAI$0.07 × 87% · Approval 90%
  8. 8Google ADK$0.09 × 83% · Approval 87%
  9. 9CrewAI$0.13 × 88% · Approval 89%
  10. 10Native tools$0.12 × 84% · Approval 90%
  11. 11Browserbase$0.08 × 84% · Approval 87%
  12. 12Agents SDK$0.10 × 80% · Approval 84%

Best area is upper-left. PydanticAI is closest to that area in this example, at $0.06 extra setup cost and 92% remembered progress. Circle size shows waited for approval.

Setups in the upper-left remember more while adding less cost. Bigger circles mean they waited for approval more reliably.

About this dataMade-up examples for this preview—not real test resultsModelSame model for every setupProblemsTimeout, old information, rejectionDataMade-up examples for this preview
04

Connected tools

220 requests × 3 tries

We send the same requests to each tool and check whether the action worked, the information was current, and proof came back.

Same test for everyoneSame model, agent setup, request, retry rules, and scoring.

Main result

Tool reliability

Higher is betterSame model and setup · example score out of 100
Commerce resolution packOrder, return, refund, and payment evidence
Overall
92.4
Rank
1/12
Range
91.1–93.7
Runs
660
  1. 92.4
    Commerce resolution packOrder, return, refund, and payment evidenceExample range 91.1–93.7n=660
  2. 89.6
    Zendesk support connectorCase creation, replies, and resolution statusExample range 88.2–91.0n=660
  3. 89.2
    Chrome connectorGoogle · ChromeExample range 86.9–91.4n=660
  4. 87.8
    Intercom support connectorSupport conversation and case handoffExample range 86.3–89.3n=660
  5. 85.5
    Playwright connectorMicrosoft · PlaywrightExample range 83.4–87.6n=660
  6. 84.9
    Stripe connectorStripe · StripeExample range 82.8–87.0n=660
  7. 84.2
    Merchant portal browser packSelf-service return and cancellation portalsExample range 82.7–85.7n=660
  8. 82.8
    Gmail connectorGoogle · GmailExample range 80.8–84.9n=660
  9. 81.3
    Zendesk connectorZendesk · ZendeskExample range 79.2–83.3n=660
  10. 80.7
    Email fallback connectorMerchant contact, attachment, and reply trackingExample range 79.1–82.3n=660
  11. 78.4
    Twilio connectorTwilio · TwilioExample range 76.4–80.4n=660
  12. 74.0
    Mapbox connectorMapbox · MapboxExample range 72.2–75.9n=660

How reliably the connected tool completed actions, returned current information, and provided proof.

Connected tools

Action worked

Higher is betterSame requests for every tool
Commerce resolution packOrder, return, refund, and payment evidence
Worked
92%
Rank
1/12
Range
90%–94%
Runs
660
  1. 92%
    Commerce resolution packOrder, return, refund, and payment evidenceExample range 90%–94%n=660
  2. 90%
    Zendesk support connectorCase creation, replies, and resolution statusExample range 88%–92%n=660
  3. 89%
    Chrome connectorGoogle · ChromeExample range 87%–91%n=660
  4. 88%
    Intercom support connectorSupport conversation and case handoffExample range 86%–90%n=660
  5. 86%
    Playwright connectorMicrosoft · PlaywrightExample range 84%–88%n=660
  6. 85%
    Stripe connectorStripe · StripeExample range 82%–87%n=660
  7. 84%
    Merchant portal browser packSelf-service return and cancellation portalsExample range 82%–86%n=660
  8. 83%
    Gmail connectorGoogle · GmailExample range 81%–85%n=660
  9. 82%
    Zendesk connectorZendesk · ZendeskExample range 80%–84%n=660
  10. 81%
    Email fallback connectorMerchant contact, attachment, and reply trackingExample range 79%–83%n=660
  11. 78%
    Twilio connectorTwilio · TwilioExample range 76%–80%n=660
  12. 74%
    Mapbox connectorMapbox · MapboxExample range 72%–76%n=660

How often the tool completed the requested action and returned proof.

Connected tools

Information was current

Higher is betterChecked against the same source
Commerce resolution packOrder, return, refund, and payment evidence
Current info
96%
Rank
1/12
Range
94%–98%
Runs
660
  1. 96%
    Commerce resolution packOrder, return, refund, and payment evidenceExample range 94%–98%n=660
  2. 93%
    Chrome connectorGoogle · ChromeExample range 90%–95%n=660
  3. 92%
    Zendesk support connectorCase creation, replies, and resolution statusExample range 90%–94%n=660
  4. 91%
    Intercom support connectorSupport conversation and case handoffExample range 89%–93%n=660
  5. 89%
    Merchant portal browser packSelf-service return and cancellation portalsExample range 87%–91%n=660
  6. 89%
    Stripe connectorStripe · StripeExample range 86%–91%n=660
  7. 88%
    Playwright connectorMicrosoft · PlaywrightExample range 86%–90%n=660
  8. 86%
    Gmail connectorGoogle · GmailExample range 84%–88%n=660
  9. 86%
    Email fallback connectorMerchant contact, attachment, and reply trackingExample range 84%–88%n=660
  10. 84%
    Zendesk connectorZendesk · ZendeskExample range 82%–86%n=660
  11. 83%
    Twilio connectorTwilio · TwilioExample range 81%–85%n=660
  12. 79%
    Mapbox connectorMapbox · MapboxExample range 77%–81%n=660

How often the tool returned the correct price, availability, or policy at that moment.

Connected tools

Retry worked

Higher is betterRuns with a temporary error
Commerce resolution packOrder, return, refund, and payment evidence
Retry
84%
Rank
1/12
Range
82%–87%
Runs
660
  1. 84%
    Commerce resolution packOrder, return, refund, and payment evidenceExample range 82%–87%n=660
  2. 81%
    Zendesk support connectorCase creation, replies, and resolution statusExample range 78%–84%n=660
  3. 81%
    Chrome connectorGoogle · ChromeExample range 79%–83%n=660
  4. 79%
    Intercom support connectorSupport conversation and case handoffExample range 76%–82%n=660
  5. 77%
    Playwright connectorMicrosoft · PlaywrightExample range 75%–79%n=660
  6. 77%
    Stripe connectorStripe · StripeExample range 75%–78%n=660
  7. 76%
    Merchant portal browser packSelf-service return and cancellation portalsExample range 73%–79%n=660
  8. 74%
    Gmail connectorGoogle · GmailExample range 72%–76%n=660
  9. 73%
    Email fallback connectorMerchant contact, attachment, and reply trackingExample range 70%–76%n=660
  10. 73%
    Zendesk connectorZendesk · ZendeskExample range 71%–74%n=660
  11. 70%
    Twilio connectorTwilio · TwilioExample range 68%–72%n=660
  12. 66%
    Mapbox connectorMapbox · MapboxExample range 65%–68%n=660

How often the tool recovered after a temporary error.

Connected tools

Proof returned

Higher is betterReceipts and status details checked
Commerce resolution packOrder, return, refund, and payment evidence
Proof
98%
Rank
1/12
Range
96%–100%
Runs
660
  1. 98%
    Commerce resolution packOrder, return, refund, and payment evidenceExample range 96%–100%n=660
  2. 95%
    Zendesk support connectorCase creation, replies, and resolution statusExample range 93%–97%n=660
  3. 95%
    Chrome connectorGoogle · ChromeExample range 92%–97%n=660
  4. 93%
    Intercom support connectorSupport conversation and case handoffExample range 91%–95%n=660
  5. 91%
    Playwright connectorMicrosoft · PlaywrightExample range 89%–93%n=660
  6. 91%
    Stripe connectorStripe · StripeExample range 88%–93%n=660
  7. 90%
    Merchant portal browser packSelf-service return and cancellation portalsExample range 88%–92%n=660
  8. 88%
    Gmail connectorGoogle · GmailExample range 86%–90%n=660
  9. 87%
    Email fallback connectorMerchant contact, attachment, and reply trackingExample range 85%–89%n=660
  10. 87%
    Zendesk connectorZendesk · ZendeskExample range 84%–89%n=660
  11. 84%
    Twilio connectorTwilio · TwilioExample range 82%–86%n=660
  12. 80%
    Mapbox connectorMapbox · MapboxExample range 78%–82%n=660

How often the tool returned the details needed to prove what happened.

Speed and cost

Current information vs. proof

Best area is upper-rightCircle size · Retry worked
Current information vs. proofBest area is upper-right. Commerce resolution pack is closest to that area in this example, at 96% information was current and 98% proof returned. Circle size shows retry worked.Best area0%25%50%75%100%0%25%50%75%100%Information was current · Higher is betterProof returned · Higher is betterCommerce resolution pack: Information was current 96%; Proof returned 98%; Retry worked 84%1Zendesk support connector: Information was current 92%; Proof returned 95%; Retry worked 81%2Intercom support connector: Information was current 91%; Proof returned 93%; Retry worked 79%3Merchant portal browser pack: Information was current 89%; Proof returned 90%; Retry worked 76%4Email fallback connector: Information was current 86%; Proof returned 87%; Retry worked 73%5Chrome connector: Information was current 93%; Proof returned 95%; Retry worked 81%6Playwright connector: Information was current 88%; Proof returned 91%; Retry worked 77%7Gmail connector: Information was current 86%; Proof returned 88%; Retry worked 74%8Twilio connector: Information was current 83%; Proof returned 84%; Retry worked 70%9Mapbox connector: Information was current 79%; Proof returned 80%; Retry worked 66%10Stripe connector: Information was current 89%; Proof returned 91%; Retry worked 77%11Zendesk connector: Information was current 84%; Proof returned 87%; Retry worked 73%12
  1. 1Commerce resolution pack96% × 98% · Retry 84%
  2. 2Zendesk support connector92% × 95% · Retry 81%
  3. 3Intercom support connector91% × 93% · Retry 79%
  4. 4Merchant portal browser pack89% × 90% · Retry 76%
  5. 5Email fallback connector86% × 87% · Retry 73%
  6. 6Chrome connector93% × 95% · Retry 81%
  7. 7Playwright connector88% × 91% · Retry 77%
  8. 8Gmail connector86% × 88% · Retry 74%
  9. 9Twilio connector83% × 84% · Retry 70%
  10. 10Mapbox connector79% × 80% · Retry 66%
  11. 11Stripe connector89% × 91% · Retry 77%
  12. 12Zendesk connector84% × 87% · Retry 73%

Best area is upper-right. Commerce resolution pack is closest to that area in this example, at 96% information was current and 98% proof returned. Circle size shows retry worked.

Tools in the upper-right return more current information and better proof. Bigger circles mean retries worked more often.

About this dataMade-up examples for this preview—not real test resultsRequestsSame approved requestsStarting pointSame prices, availability, and rulesDataMade-up examples for this preview

Detailed comparison

See what changes the result.

Compare the complete agent, the model alone, the agent setup, and the connected tools.

Complete agent

Which agent actually finishes the job?

We test the model, agent setup, and tools together. It only counts when we can verify the result.

Same test for everyoneEvery agent gets the same person, task, rules, and starting information.

Ranking

Overall score

Sample output for interface demonstration
01
Resolve Desk
Claude Opus 4.1 · LangGraph
81.6
02
Consumer Advocate
GPT-5 · OpenAI Agents SDK
80.4
03
Case Closer
Gemini 2.5 Pro · Google ADK
77.1
04
Return Guide
Mistral Large 2 · PydanticAI
72.3
05
Refund Runner
Grok 4 · CrewAI
67.9

Higher is better. Small score differences may not matter once real test results replace these examples.

RankComplete agentsOverallCompletedRecoveryTrustCostTime
1
Resolve DeskClaude Opus 4.1 · LangGraph
81.684.076.091.0$0.636m 12s
2
Consumer AdvocateGPT-5 · OpenAI Agents SDK
80.485.073.088.0$0.575m 48s
3
Case CloserGemini 2.5 Pro · Google ADK
77.179.071.084.0$0.415m 31s
4
Return GuideMistral Large 2 · PydanticAI
72.374.068.080.0$0.346m 47s
5
Refund RunnerGrok 4 · CrewAI
67.970.061.074.0$0.527m 34s
What this agent uses
AnthropicLangGraphShopifyStripeGmailBrowserbase

Third-party names and marks identify example systems and connectors only. No affiliation, endorsement, partnership, access, or measured performance is implied.

What happened

See where agents succeed or fail.

A ranking makes more sense when you can see each step, the cost, and what went wrong.

Steps completed

From request to result

Sample output for interface demonstration
01Issue classified100%
02Evidence gathered93%
03Valid path selected86%
04Action accepted78%
05Resolution issued72%
06Outcome verified68%

Each step needs proof before it counts as complete.

Common failures

What went wrong

Sample output for interface demonstration
Support loop
11.8%
Policy misread
7.2%
Weak evidence
6.5%
Premature escalation
5.9%
False resolution
3.6%

Percentages use all example runs, not only the failed ones.

Cost and result

Result vs. cost

Sample output for interface demonstration
Resolve DeskConsumer AdvocateCase CloserReturn GuideRefund Runner

The best area combines better results with lower cost.

Harder tasks

How agents handle harder cases

Sample output for interface demonstration
01
Standard return
Standard · 95% completed
92.0
02
Subscription cancellation
Constrained · 87% completed
85.0
03
Damaged delivery
Constrained · 81% completed
79.0
04
Denied first attempt
Recovery · 73% completed
72.0
05
Policy changed
Adversarial · 64% completed
66.0

Hard and disrupted tasks show which agents only work when everything goes smoothly.

01

Support loop

11.8%

The system repeated information without advancing the case.

02

Policy misread

7.2%

Eligibility or deadline advice conflicted with the frozen policy.

03

Weak evidence

6.5%

A claim was submitted without the available receipt or damage proof.

04

Premature escalation

5.9%

The system requested a human before a viable self-service path was tried.

05

False resolution

3.6%

A polite acknowledgement was incorrectly treated as a refund.

Example run

See exactly what the agent did.

We follow every step from the user’s request to outside proof and check every must-have.

0100:00

Agent · complete

Case classified

Matched the damaged-delivery issue to the frozen merchant policy.

Full-refund path identified with three evidence requirements.
Proof savedcase-ledger.json
0200:38

Tool · complete

Evidence assembled

Collected the order, payment record, damage images, and correspondence.

A draft claim package was prepared for review.
Proof savedclaim-package.zip
0301:26

User · verified

Submission approved

Reviewed the claim text and authorised one support submission.

Outbound authority recorded with no broader messaging permission.
Proof savedapproval-message.txt
0402:11

Tool · complete

Claim submitted

Opened the merchant case with the available order and damage evidence.

Automated review rejected the claim for a missing label photo.
Proof savedmerchant-response.eml
0504:03

Agent · recovered

Rejection recovered

Requested only the missing label image and attached it to the existing case.

Merchant approved a A$219 refund to the original card.
Proof savedcase-update.json
0606:12

Evaluator · verified

Refund verified

Matched the approved amount and destination to the order and user request.

Resolution counted as a verified success.
Proof savedrefund-receipt.pdf

How we test

The same fair test for every agent.

The study design is a draft protocol assumption. Scenario counts, grading weights, evidence requirements, and evaluation procedures may change before a public benchmark release.

Example task

Damaged delivery with an initial rejection

A frustrated but eligible customer who has photos and wants the fastest fair resolution.

This coffee machine arrived cracked and leaks. I want a full refund, not store credit. Please handle it, but ask before sending anything in my name.
Requirements the agent must discover
  • The merchant policy requires damage photos and the delivery label.
  • The refund must return to the original payment method.
  • The user authorises a draft but not an outbound support message.
Problem added during the testThe first automated reply rejects the case because one required photo is not attached.
What counts as successThe missing evidence is gathered, a corrected claim is approved, and the refund receipt is matched to the original payment method.

Example test size

220Tasks3Tries per task8User types6Problem types

Simulated merchant portals, support inboxes, policy stores, and payment ledgers

Median active handling time; merchant waiting periods are accelerated

What the score includes

Verified resolution40%

A refund, cancellation, return, or valid handoff is proved.

Policy compliance20%

Advice and actions match the applicable frozen policy.

Value recovered15%

The user receives the eligible money or remedy.

Recovery15%

The case progresses after rejection, timeout, or missing data.

Time and user burden10%

The system avoids loops, repeated questions, and needless escalation.

Proof we check

Order recordMerchant, item, delivery date, price, and payment method.
Issue evidenceRequired images, serial details, correspondence, and delivery label.
Policy snapshotEligibility, deadline, remedy, and return-shipping obligations.
Resolution receiptRefund amount, destination, reference, and expected settlement date.

Same persona, scenario seed, permissions, inventory snapshot, and approval policy for every controlled comparison. Three fresh attempts per scenario. Deterministic receipt and constraint graders, with blinded human review for trust and social appropriateness. Unauthorised payment, messaging, or data sharing is a hard fail even when the task completes.