Agent tests/DIN-02/Preview

Real-world task test · DIN-02

Group Dinner Coordination

Can an agent turn six people’s preferences into a dinner everyone agrees to?

We test whether the agent collects everyone’s availability and needs, finds a fair option, books it, keeps the group informed, and handles last-minute changes.

Highlights · DIN-02

Made-up example data
01

Complete agents

180 tasks × 3 tries = 540 example runs

We test the model, agent setup, and tools together. It only counts when we can verify the result.

Same test for everyoneEvery agent gets the same person, task, rules, and starting information.

Main result

Dinners booked

Higher is betterBased on the example runs
Gather AgentGPT-5 · OpenAI Agents SDK
Booked
78%
Rank
1/12
Range
75%–81%
Runs
540
  1. 78%
    Gather AgentGPT-5 · OpenAI Agents SDKExample range 75%–81%n=540
  2. 76%
    Table CaptainClaude Opus 4.1 · LangGraphExample range 74%–78%n=540
  3. 74%
    Scout AIAnthropic · LangGraphExample range 72%–76%n=540
  4. 73%
    Plan TogetherGemini 2.5 Pro · Google ADKExample range 70%–76%n=540
  5. 73%
    Compass AgentOpenAI · Agents SDKExample range 71%–75%n=540
  6. 70%
    Waypoint ProxAI · Native toolsExample range 68%–71%n=540
  7. 69%
    Nova AssistDeepSeek · LangChainExample range 67%–70%n=540
  8. 68%
    Relay ConciergeGoogle · Google ADKExample range 66%–70%n=540
  9. 68%
    Supper CircleMistral Large 2 · PydanticAIExample range 65%–71%n=540
  10. 65%
    Night OutGrok 4 · CrewAIExample range 62%–68%n=540
  11. 62%
    Orbit PlannerMistral · PydanticAIExample range 61%–64%n=540
  12. 58%
    Task PilotMeta · CrewAIExample range 57%–60%n=540

The group approved the restaurant and time, and the agent booked a real table.

Cost

Cheapest booking

Lower is betterExample cost in USD
Supper CircleMistral Large 2 · PydanticAI
Cost
$0.32
Rank
1/12
Range
$0.27–$0.37
Runs
540
  1. $0.32
    Supper CircleMistral Large 2 · PydanticAIExample range $0.27–$0.37n=540
  2. $0.39
    Plan TogetherGemini 2.5 Pro · Google ADKExample range $0.35–$0.43n=540
  3. $0.40
    Orbit PlannerMistral · PydanticAIExample range $0.37–$0.43n=540
  4. $0.47
    Relay ConciergeGoogle · Google ADKExample range $0.44–$0.50n=540
  5. $0.49
    Night OutGrok 4 · CrewAIExample range $0.44–$0.54n=540
  6. $0.54
    Gather AgentGPT-5 · OpenAI Agents SDKExample range $0.50–$0.58n=540
  7. $0.58
    Table CaptainClaude Opus 4.1 · LangGraphExample range $0.55–$0.61n=540
  8. $0.63
    Scout AIAnthropic · LangGraphExample range $0.60–$0.66n=540
  9. $0.64
    Task PilotMeta · CrewAIExample range $0.61–$0.67n=540
  10. $0.65
    Compass AgentOpenAI · Agents SDKExample range $0.62–$0.68n=540
  11. $0.75
    Waypoint ProxAI · Native toolsExample range $0.72–$0.78n=540
  12. $0.78
    Nova AssistDeepSeek · LangChainExample range $0.75–$0.81n=540

The estimated model and tool cost for each completed task.

Speed

Fastest to book

Lower is betterLower is better
Plan TogetherGemini 2.5 Pro · Google ADK
Time
8m 50s
Rank
1/12
Range
8m 32s–9m 08s
Runs
540
  1. 8m 50s
    Plan TogetherGemini 2.5 Pro · Google ADKExample range 8m 32s–9m 08sn=540
  2. 9m 32s
    Table CaptainClaude Opus 4.1 · LangGraphExample range 9m 18s–9m 46sn=540
  3. 10m 18s
    Gather AgentGPT-5 · OpenAI Agents SDKExample range 10m 02s–10m 34sn=540
  4. 10m 44s
    Compass AgentOpenAI · Agents SDKExample range 10m 30s–10m 58sn=540
  5. 10m 44s
    Relay ConciergeGoogle · Google ADKExample range 10m 30s–10m 58sn=540
  6. 11m 24s
    Supper CircleMistral Large 2 · PydanticAIExample range 11m 04s–11m 44sn=540
  7. 12m 03s
    Scout AIAnthropic · LangGraphExample range 11m 49s–12m 17sn=540
  8. 12m 22s
    Night OutGrok 4 · CrewAIExample range 12m 00s–12m 44sn=540
  9. 12m 52s
    Nova AssistDeepSeek · LangChainExample range 12m 38s–13m 06sn=540
  10. 14m 22s
    Orbit PlannerMistral · PydanticAIExample range 14m 08s–14m 36sn=540
  11. 14m 22s
    Waypoint ProxAI · Native toolsExample range 14m 08s–14m 36sn=540
  12. 16m 08s
    Task PilotMeta · CrewAIExample range 15m 54s–16m 22sn=540

How long the agent actively worked on the task.

Complete agents

Group agreed and attended

Higher is betterAll group tasks · example range
Table CaptainClaude Opus 4.1 · LangGraph
Agreed + went
64.2%
Rank
1/12
Range
62.0%–66.4%
Runs
540
  1. 64.2%
    Table CaptainClaude Opus 4.1 · LangGraphExample range 62.0%–66.4%n=540
  2. 62.8%
    Gather AgentGPT-5 · OpenAI Agents SDKExample range 60.6%–65.0%n=540
  3. 61.0%
    Compass AgentOpenAI · Agents SDKExample range 59.4%–62.5%n=540
  4. 60.6%
    Plan TogetherGemini 2.5 Pro · Google ADKExample range 58.4%–62.8%n=540
  5. 58.7%
    Scout AIAnthropic · LangGraphExample range 57.2%–60.2%n=540
  6. 56.7%
    Nova AssistDeepSeek · LangChainExample range 55.3%–58.1%n=540
  7. 55.7%
    Relay ConciergeGoogle · Google ADKExample range 54.3%–57.1%n=540
  8. 54.4%
    Waypoint ProxAI · Native toolsExample range 53.0%–55.8%n=540
  9. 54.1%
    Supper CircleMistral Large 2 · PydanticAIExample range 51.9%–56.3%n=540
  10. 49.7%
    Night OutGrok 4 · CrewAIExample range 47.5%–51.9%n=540
  11. 48.3%
    Orbit PlannerMistral · PydanticAIExample range 46.9%–49.7%n=540
  12. 43.1%
    Task PilotMeta · CrewAIExample range 41.7%–44.5%n=540

At least 80% of the group agreed to the plan, everyone’s must-haves were met, and at least 80% attended.

Complete agents

Overall score

Higher is betterExample score out of 100
Table CaptainClaude Opus 4.1 · LangGraph
Overall
74.8
Rank
1/12
Range
73.1–76.5
Runs
540
  1. 74.8
    Table CaptainClaude Opus 4.1 · LangGraphExample range 73.1–76.5n=540
  2. 73.5
    Gather AgentGPT-5 · OpenAI Agents SDKExample range 71.7–75.3n=540
  3. 71.5
    Compass AgentOpenAI · Agents SDKExample range 69.8–73.3n=540
  4. 71.2
    Plan TogetherGemini 2.5 Pro · Google ADKExample range 69.3–73.1n=540
  5. 69.4
    Scout AIAnthropic · LangGraphExample range 67.7–71.1n=540
  6. 67.3
    Nova AssistDeepSeek · LangChainExample range 65.6–69.0n=540
  7. 66.4
    Supper CircleMistral Large 2 · PydanticAIExample range 64.5–68.3n=540
  8. 66.3
    Relay ConciergeGoogle · Google ADKExample range 64.6–67.9n=540
  9. 65.2
    Waypoint ProxAI · Native toolsExample range 63.5–66.8n=540
  10. 63.1
    Night OutGrok 4 · CrewAIExample range 61.1–65.1n=540
  11. 60.6
    Orbit PlannerMistral · PydanticAIExample range 59.1–62.1n=540
  12. 56.5
    Task PilotMeta · CrewAIExample range 55.0–57.9n=540

A simple score combining task success, following requirements, recovery, cost, and trust.

Complete agents

Rebooked after a change

Higher is betterRuns where something went wrong
Table CaptainClaude Opus 4.1 · LangGraph
Recovery
69%
Rank
1/12
Range
66%–72%
Runs
540
  1. 69%
    Table CaptainClaude Opus 4.1 · LangGraphExample range 66%–72%n=540
  2. 66%
    Gather AgentGPT-5 · OpenAI Agents SDKExample range 63%–69%n=540
  3. 66%
    Compass AgentOpenAI · Agents SDKExample range 64%–67%n=540
  4. 65%
    Plan TogetherGemini 2.5 Pro · Google ADKExample range 62%–68%n=540
  5. 62%
    Scout AIAnthropic · LangGraphExample range 60%–63%n=540
  6. 62%
    Nova AssistDeepSeek · LangChainExample range 60%–63%n=540
  7. 61%
    Supper CircleMistral Large 2 · PydanticAIExample range 58%–64%n=540
  8. 60%
    Relay ConciergeGoogle · Google ADKExample range 59%–62%n=540
  9. 58%
    Waypoint ProxAI · Native toolsExample range 56%–59%n=540
  10. 57%
    Night OutGrok 4 · CrewAIExample range 54%–60%n=540
  11. 55%
    Orbit PlannerMistral · PydanticAIExample range 54%–57%n=540
  12. 50%
    Task PilotMeta · CrewAIExample range 49%–52%n=540

How often the agent still finished after a price change, tool error, rejection, or timeout.

Complete agents

Followed approval rules

Higher is betterExample score out of 100
Table CaptainClaude Opus 4.1 · LangGraph
Trust
88%
Rank
1/12
Range
86%–90%
Runs
540
  1. 88%
    Table CaptainClaude Opus 4.1 · LangGraphExample range 86%–90%n=540
  2. 85%
    Gather AgentGPT-5 · OpenAI Agents SDKExample range 83%–87%n=540
  3. 85%
    Compass AgentOpenAI · Agents SDKExample range 83%–87%n=540
  4. 82%
    Plan TogetherGemini 2.5 Pro · Google ADKExample range 80%–84%n=540
  5. 81%
    Scout AIAnthropic · LangGraphExample range 79%–83%n=540
  6. 81%
    Nova AssistDeepSeek · LangChainExample range 78%–83%n=540
  7. 78%
    Supper CircleMistral Large 2 · PydanticAIExample range 76%–80%n=540
  8. 77%
    Relay ConciergeGoogle · Google ADKExample range 75%–79%n=540
  9. 77%
    Waypoint ProxAI · Native toolsExample range 75%–79%n=540
  10. 73%
    Night OutGrok 4 · CrewAIExample range 71%–75%n=540
  11. 72%
    Orbit PlannerMistral · PydanticAIExample range 70%–74%n=540
  12. 66%
    Task PilotMeta · CrewAIExample range 65%–68%n=540

Whether the agent asked before important actions, showed its proof, and was honest about uncertainty.

Speed and cost

Time vs. cost

Best area is lower-leftCircle size · Dinners booked
Time vs. costBest area is lower-left. Plan Together is closest to that area in this example, at 8m 50s fastest to book and $0.39 cheapest booking. Circle size shows dinners booked.Best area7m 49s10m 09s12m 29s14m 49s17m 10s$0.26$0.40$0.55$0.70$0.85Fastest to book · Lower is betterCheapest booking · Lower is betterTable Captain: Fastest to book 9m 32s; Cheapest booking $0.58; Dinners booked 76%1Gather Agent: Fastest to book 10m 18s; Cheapest booking $0.54; Dinners booked 78%2Plan Together: Fastest to book 8m 50s; Cheapest booking $0.39; Dinners booked 73%3Supper Circle: Fastest to book 11m 24s; Cheapest booking $0.32; Dinners booked 68%4Night Out: Fastest to book 12m 22s; Cheapest booking $0.49; Dinners booked 65%5Compass Agent: Fastest to book 10m 44s; Cheapest booking $0.65; Dinners booked 73%6Scout AI: Fastest to book 12m 03s; Cheapest booking $0.63; Dinners booked 74%7Relay Concierge: Fastest to book 10m 44s; Cheapest booking $0.47; Dinners booked 68%8Orbit Planner: Fastest to book 14m 22s; Cheapest booking $0.40; Dinners booked 62%9Task Pilot: Fastest to book 16m 08s; Cheapest booking $0.64; Dinners booked 58%10Nova Assist: Fastest to book 12m 52s; Cheapest booking $0.78; Dinners booked 69%11Waypoint Pro: Fastest to book 14m 22s; Cheapest booking $0.75; Dinners booked 70%12
  1. 1Table Captain9m 32s × $0.58 · Booked 76%
  2. 2Gather Agent10m 18s × $0.54 · Booked 78%
  3. 3Plan Together8m 50s × $0.39 · Booked 73%
  4. 4Supper Circle11m 24s × $0.32 · Booked 68%
  5. 5Night Out12m 22s × $0.49 · Booked 65%
  6. 6Compass Agent10m 44s × $0.65 · Booked 73%
  7. 7Scout AI12m 03s × $0.63 · Booked 74%
  8. 8Relay Concierge10m 44s × $0.47 · Booked 68%
  9. 9Orbit Planner14m 22s × $0.40 · Booked 62%
  10. 10Task Pilot16m 08s × $0.64 · Booked 58%
  11. 11Nova Assist12m 52s × $0.78 · Booked 69%
  12. 12Waypoint Pro14m 22s × $0.75 · Booked 70%

Best area is lower-left. Plan Together is closest to that area in this example, at 8m 50s fastest to book and $0.39 cheapest booking. Circle size shows dinners booked.

Agents closer to the lower-left finished faster and cost less. Bigger circles booked more dinners.

About this dataMade-up examples for this preview—not real test resultsRangeGrouped by dinner taskProof requiredReceipt or other outside proofDataMade-up examples for this preview
02

Models only

180 tasks × 3 tries

Each model sees the same task and options. Tools are turned off, so the model cannot take actions.

Same test for everyoneSame task, options, and scoring. Tools are off.

Main result

Overall reasoning

Higher is betterTools off · example score out of 100
Claude Opus 4.1Anthropic · tools disabled
Overall
86.2
Rank
1/12
Range
84.8–87.6
Runs
540
  1. 86.2
    Claude Opus 4.1Anthropic · tools disabledExample range 84.8–87.6n=540
  2. 84.9
    GPT-5OpenAI · tools disabledExample range 83.4–86.4n=540
  3. 83.0
    GPT-4.1OpenAI · tools disabledExample range 80.9–85.0n=540
  4. 81.7
    Gemini 2.5 ProGoogle · tools disabledExample range 80.1–83.3n=540
  5. 80.8
    Claude 3.7 SonnetAnthropic · tools disabledExample range 78.8–82.8n=540
  6. 78.7
    DeepSeek V3DeepSeek · tools disabledExample range 76.7–80.7n=540
  7. 76.8
    Gemini 2.0 FlashGoogle · tools disabledExample range 74.8–78.7n=540
  8. 76.6
    Grok 3xAI · tools disabledExample range 74.6–78.5n=540
  9. 75.4
    Mistral Large 2Mistral AI · tools disabledExample range 73.8–77.0n=540
  10. 72.8
    Grok 4xAI · tools disabledExample range 71.1–74.5n=540
  11. 69.6
    Mistral LargeMistral · tools disabledExample range 67.9–71.3n=540
  12. 66.1
    Llama 4 MaverickMeta · tools disabledExample range 64.5–67.8n=540

How well the model understood the task and made a choice. The model could not use tools or take actions.

Models only

Remembered requirements

Higher is betterSame task information for every model
Claude Opus 4.1Anthropic · tools disabled
Requirements
93%
Rank
1/12
Range
91%–95%
Runs
540
  1. 93%
    Claude Opus 4.1Anthropic · tools disabledExample range 91%–95%n=540
  2. 91%
    GPT-5OpenAI · tools disabledExample range 89%–93%n=540
  3. 90%
    GPT-4.1OpenAI · tools disabledExample range 88%–92%n=540
  4. 89%
    Gemini 2.5 ProGoogle · tools disabledExample range 87%–91%n=540
  5. 87%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 85%–89%n=540
  6. 86%
    DeepSeek V3DeepSeek · tools disabledExample range 83%–88%n=540
  7. 84%
    Gemini 2.0 FlashGoogle · tools disabledExample range 82%–86%n=540
  8. 84%
    Mistral Large 2Mistral AI · tools disabledExample range 82%–86%n=540
  9. 83%
    Grok 3xAI · tools disabledExample range 81%–85%n=540
  10. 80%
    Grok 4xAI · tools disabledExample range 78%–82%n=540
  11. 78%
    Mistral LargeMistral · tools disabledExample range 76%–80%n=540
  12. 73%
    Llama 4 MaverickMeta · tools disabledExample range 72%–75%n=540

How often the model remembered the user’s must-haves when making a choice.

Models only

Choice quality

Higher is betterSame options and scoring
GPT-5OpenAI · tools disabled
Choice
88%
Rank
1/12
Range
86%–90%
Runs
540
  1. 88%
    GPT-5OpenAI · tools disabledExample range 86%–90%n=540
  2. 87%
    Claude Opus 4.1Anthropic · tools disabledExample range 85%–89%n=540
  3. 84%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 82%–86%n=540
  4. 84%
    GPT-4.1OpenAI · tools disabledExample range 82%–86%n=540
  5. 83%
    Gemini 2.5 ProGoogle · tools disabledExample range 81%–85%n=540
  6. 80%
    Grok 3xAI · tools disabledExample range 78%–82%n=540
  7. 80%
    DeepSeek V3DeepSeek · tools disabledExample range 78%–81%n=540
  8. 78%
    Gemini 2.0 FlashGoogle · tools disabledExample range 76%–80%n=540
  9. 76%
    Mistral Large 2Mistral AI · tools disabledExample range 74%–78%n=540
  10. 75%
    Grok 4xAI · tools disabledExample range 73%–77%n=540
  11. 70%
    Mistral LargeMistral · tools disabledExample range 68%–72%n=540
  12. 68%
    Llama 4 MaverickMeta · tools disabledExample range 67%–70%n=540

How good the model’s choice was when every model saw the same options.

Models only

Action plan works

Higher is betterPlans reviewed without running them
GPT-5OpenAI · tools disabled
Plan
84%
Rank
1/12
Range
82%–86%
Runs
540
  1. 84%
    GPT-5OpenAI · tools disabledExample range 82%–86%n=540
  2. 83%
    Claude Opus 4.1Anthropic · tools disabledExample range 81%–85%n=540
  3. 81%
    Gemini 2.5 ProGoogle · tools disabledExample range 79%–83%n=540
  4. 80%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 78%–82%n=540
  5. 80%
    GPT-4.1OpenAI · tools disabledExample range 78%–82%n=540
  6. 76%
    Gemini 2.0 FlashGoogle · tools disabledExample range 74%–78%n=540
  7. 76%
    Grok 3xAI · tools disabledExample range 74%–78%n=540
  8. 76%
    DeepSeek V3DeepSeek · tools disabledExample range 74%–77%n=540
  9. 74%
    Mistral Large 2Mistral AI · tools disabledExample range 72%–77%n=540
  10. 72%
    Grok 4xAI · tools disabledExample range 69%–75%n=540
  11. 68%
    Mistral LargeMistral · tools disabledExample range 66%–70%n=540
  12. 65%
    Llama 4 MaverickMeta · tools disabledExample range 64%–67%n=540

How often the proposed steps were possible, in the right order, and safe to follow.

Models only

Honest about uncertainty

Higher is betterSame scoring for every model
Claude Opus 4.1Anthropic · tools disabled
Honesty
90%
Rank
1/12
Range
88%–92%
Runs
540
  1. 90%
    Claude Opus 4.1Anthropic · tools disabledExample range 88%–92%n=540
  2. 87%
    GPT-4.1OpenAI · tools disabledExample range 85%–89%n=540
  3. 86%
    GPT-5OpenAI · tools disabledExample range 84%–88%n=540
  4. 83%
    Gemini 2.5 ProGoogle · tools disabledExample range 81%–85%n=540
  5. 83%
    DeepSeek V3DeepSeek · tools disabledExample range 80%–85%n=540
  6. 82%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 80%–84%n=540
  7. 79%
    Mistral Large 2Mistral AI · tools disabledExample range 77%–81%n=540
  8. 78%
    Gemini 2.0 FlashGoogle · tools disabledExample range 76%–80%n=540
  9. 78%
    Grok 3xAI · tools disabledExample range 76%–80%n=540
  10. 73%
    Mistral LargeMistral · tools disabledExample range 71%–75%n=540
  11. 72%
    Grok 4xAI · tools disabledExample range 70%–74%n=540
  12. 65%
    Llama 4 MaverickMeta · tools disabledExample range 64%–67%n=540

Whether the model used the available proof and avoided claiming success when it was unsure.

Speed and cost

Good choice vs. workable plan

Best area is upper-rightCircle size · Honest about uncertainty
Good choice vs. workable planBest area is upper-right. GPT-5 is closest to that area in this example, at 84% action plan works and 88% choice quality. Circle size shows honest about uncertainty.Best area0%25%50%75%100%0%25%50%75%100%Action plan works · Higher is betterChoice quality · Higher is betterClaude Opus 4.1: Action plan works 83%; Choice quality 87%; Honest about uncertainty 90%1GPT-5: Action plan works 84%; Choice quality 88%; Honest about uncertainty 86%2Gemini 2.5 Pro: Action plan works 81%; Choice quality 83%; Honest about uncertainty 83%3Mistral Large 2: Action plan works 74%; Choice quality 76%; Honest about uncertainty 79%4Grok 4: Action plan works 72%; Choice quality 75%; Honest about uncertainty 72%5GPT-4.1: Action plan works 80%; Choice quality 84%; Honest about uncertainty 87%6Claude 3.7 Sonnet: Action plan works 80%; Choice quality 84%; Honest about uncertainty 82%7Gemini 2.0 Flash: Action plan works 76%; Choice quality 78%; Honest about uncertainty 78%8Mistral Large: Action plan works 68%; Choice quality 70%; Honest about uncertainty 73%9Llama 4 Maverick: Action plan works 65%; Choice quality 68%; Honest about uncertainty 65%10DeepSeek V3: Action plan works 76%; Choice quality 80%; Honest about uncertainty 83%11Grok 3: Action plan works 76%; Choice quality 80%; Honest about uncertainty 78%12
  1. 1Claude Opus 4.183% × 87% · Honesty 90%
  2. 2GPT-584% × 88% · Honesty 86%
  3. 3Gemini 2.5 Pro81% × 83% · Honesty 83%
  4. 4Mistral Large 274% × 76% · Honesty 79%
  5. 5Grok 472% × 75% · Honesty 72%
  6. 6GPT-4.180% × 84% · Honesty 87%
  7. 7Claude 3.7 Sonnet80% × 84% · Honesty 82%
  8. 8Gemini 2.0 Flash76% × 78% · Honesty 78%
  9. 9Mistral Large68% × 70% · Honesty 73%
  10. 10Llama 4 Maverick65% × 68% · Honesty 65%
  11. 11DeepSeek V376% × 80% · Honesty 83%
  12. 12Grok 376% × 80% · Honesty 78%

Best area is upper-right. GPT-5 is closest to that area in this example, at 84% action plan works and 88% choice quality. Circle size shows honest about uncertainty.

Models in the upper-right make better choices and produce plans that are more likely to work. Bigger circles mean the model was more honest about uncertainty.

About this dataMade-up examples for this preview—not real test resultsActionsTurned offOptionsSame options for every modelDataMade-up examples for this preview
03

Agent setup

180 tasks × 6 problem types

We keep the model and tools the same and change only the workflow that runs them.

Same test for everyoneSame model, tools, task, and problems.

Main result

Agent setup score

Higher is betterSame model and tools · example score out of 100
LangGraphLangChain · frozen model and adapters
Overall
84.6
Rank
1/12
Range
83.1–86.1
Runs
540
  1. 84.6
    LangGraphLangChain · frozen model and adaptersExample range 83.1–86.1n=540
  2. 82.9
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 81.3–84.5n=540
  3. 81.3
    Agents SDKOpenAI · Agents SDKExample range 79.3–83.4n=540
  4. 79.3
    Google ADKGoogle · frozen model and adaptersExample range 77.6–81.0n=540
  5. 78.8
    LangGraphLangChain · LangGraphExample range 76.8–80.8n=540
  6. 77.9
    CrewAICrewAI · CrewAIExample range 76.0–79.9n=540
  7. 77.1
    PydanticAIPydantic · frozen model and adaptersExample range 75.4–78.8n=540
  8. 75.4
    Native toolsAnthropic · Native toolsExample range 73.5–77.3n=540
  9. 74.3
    PydanticAIPydantic · PydanticAIExample range 72.5–76.2n=540
  10. 71.3
    Google ADKGoogle · Google ADKExample range 69.5–73.1n=540
  11. 71.0
    BrowserbaseBrowserbase · BrowserbaseExample range 69.2–72.7n=540
  12. 67.9
    Agents SDKOpenAI · Agents SDKExample range 66.2–69.6n=540

How well the workflow ran the task when the model and tools stayed the same.

Agent setup

Ran the plan correctly

Higher is betterSame approved plan
OpenAI Agents SDKOpenAI · frozen model and adapters
Execution
83%
Rank
1/12
Range
81%–85%
Runs
540
  1. 83%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 81%–85%n=540
  2. 81%
    LangGraphLangChain · frozen model and adaptersExample range 79%–83%n=540
  3. 79%
    Google ADKGoogle · frozen model and adaptersExample range 77%–81%n=540
  4. 79%
    LangGraphLangChain · LangGraphExample range 77%–81%n=540
  5. 78%
    Agents SDKOpenAI · Agents SDKExample range 76%–80%n=540
  6. 76%
    Native toolsAnthropic · Native toolsExample range 74%–77%n=540
  7. 75%
    PydanticAIPydantic · frozen model and adaptersExample range 73%–77%n=540
  8. 74%
    CrewAICrewAI · CrewAIExample range 72%–76%n=540
  9. 74%
    PydanticAIPydantic · PydanticAIExample range 72%–76%n=540
  10. 71%
    BrowserbaseBrowserbase · BrowserbaseExample range 69%–72%n=540
  11. 69%
    Google ADKGoogle · Google ADKExample range 67%–71%n=540
  12. 66%
    Agents SDKOpenAI · Agents SDKExample range 64%–67%n=540

How often the setup followed the approved steps in the right order without losing progress.

Agent setup

Recovered from an error

Higher is betterRuns where something went wrong
LangGraphLangChain · frozen model and adapters
Recovery
78%
Rank
1/12
Range
75%–81%
Runs
540
  1. 78%
    LangGraphLangChain · frozen model and adaptersExample range 75%–81%n=540
  2. 75%
    Agents SDKOpenAI · Agents SDKExample range 73%–77%n=540
  3. 73%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 70%–76%n=540
  4. 72%
    Google ADKGoogle · frozen model and adaptersExample range 69%–75%n=540
  5. 71%
    CrewAICrewAI · CrewAIExample range 70%–73%n=540
  6. 71%
    PydanticAIPydantic · frozen model and adaptersExample range 68%–74%n=540
  7. 69%
    LangGraphLangChain · LangGraphExample range 67%–71%n=540
  8. 67%
    PydanticAIPydantic · PydanticAIExample range 65%–69%n=540
  9. 66%
    Native toolsAnthropic · Native toolsExample range 64%–67%n=540
  10. 65%
    Google ADKGoogle · Google ADKExample range 64%–67%n=540
  11. 64%
    BrowserbaseBrowserbase · BrowserbaseExample range 62%–65%n=540
  12. 62%
    Agents SDKOpenAI · Agents SDKExample range 60%–63%n=540

How often the setup recovered from a timeout, outdated information, price change, or rejection.

Agent setup

Remembered progress

Higher is betterProgress record checked
LangGraphLangChain · frozen model and adapters
Memory
94%
Rank
1/12
Range
92%–96%
Runs
540
  1. 94%
    LangGraphLangChain · frozen model and adaptersExample range 92%–96%n=540
  2. 91%
    Agents SDKOpenAI · Agents SDKExample range 88%–93%n=540
  3. 89%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 87%–91%n=540
  4. 88%
    PydanticAIPydantic · frozen model and adaptersExample range 86%–90%n=540
  5. 87%
    CrewAICrewAI · CrewAIExample range 85%–90%n=540
  6. 87%
    Google ADKGoogle · frozen model and adaptersExample range 85%–89%n=540
  7. 85%
    LangGraphLangChain · LangGraphExample range 83%–87%n=540
  8. 82%
    Google ADKGoogle · Google ADKExample range 80%–84%n=540
  9. 82%
    PydanticAIPydantic · PydanticAIExample range 80%–84%n=540
  10. 82%
    Native toolsAnthropic · Native toolsExample range 79%–84%n=540
  11. 79%
    Agents SDKOpenAI · Agents SDKExample range 77%–81%n=540
  12. 79%
    BrowserbaseBrowserbase · BrowserbaseExample range 77%–81%n=540

Whether the setup remembered requirements, approvals, evidence, and completed steps.

Agent setup

Waited for approval

Higher is betterApproval checks
LangGraphLangChain · frozen model and adapters
Approval
95%
Rank
1/12
Range
93%–97%
Runs
540
  1. 95%
    LangGraphLangChain · frozen model and adaptersExample range 93%–97%n=540
  2. 94%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 92%–96%n=540
  3. 92%
    PydanticAIPydantic · frozen model and adaptersExample range 90%–94%n=540
  4. 92%
    Agents SDKOpenAI · Agents SDKExample range 89%–94%n=540
  5. 91%
    Google ADKGoogle · frozen model and adaptersExample range 89%–93%n=540
  6. 90%
    LangGraphLangChain · LangGraphExample range 88%–92%n=540
  7. 88%
    CrewAICrewAI · CrewAIExample range 86%–91%n=540
  8. 87%
    Native toolsAnthropic · Native toolsExample range 84%–89%n=540
  9. 86%
    Google ADKGoogle · Google ADKExample range 84%–88%n=540
  10. 86%
    PydanticAIPydantic · PydanticAIExample range 84%–88%n=540
  11. 83%
    Agents SDKOpenAI · Agents SDKExample range 81%–85%n=540
  12. 83%
    BrowserbaseBrowserbase · BrowserbaseExample range 81%–85%n=540

How often the setup asked the user before an important or hard-to-reverse action.

Cost

Extra setup cost

Lower is betterExample USD above the baseline
PydanticAIPydantic · frozen model and adapters
Extra cost
$0.06
Rank
1/12
Range
$0.03–$0.09
Runs
540
  1. $0.06
    PydanticAIPydantic · frozen model and adaptersExample range $0.03–$0.09n=540
  2. $0.07
    Google ADKGoogle · frozen model and adaptersExample range $0.04–$0.10n=540
  3. $0.08
    Google ADKGoogle · Google ADKExample range $0.05–$0.11n=540
  4. $0.09
    PydanticAIPydantic · PydanticAIExample range $0.06–$0.12n=540
  5. $0.09
    Agents SDKOpenAI · Agents SDKExample range $0.06–$0.12n=540
  6. $0.10
    BrowserbaseBrowserbase · BrowserbaseExample range $0.07–$0.13n=540
  7. $0.10
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range $0.08–$0.12n=540
  8. $0.12
    LangGraphLangChain · LangGraphExample range $0.09–$0.15n=540
  9. $0.12
    LangGraphLangChain · frozen model and adaptersExample range $0.10–$0.14n=540
  10. $0.14
    Agents SDKOpenAI · Agents SDKExample range $0.11–$0.17n=540
  11. $0.14
    Native toolsAnthropic · Native toolsExample range $0.11–$0.17n=540
  12. $0.16
    CrewAICrewAI · CrewAIExample range $0.13–$0.19n=540

The extra cost added by the agent setup for each completed task.

Speed and cost

Memory vs. extra cost

Best area is upper-leftCircle size · Waited for approval
Memory vs. extra costBest area is upper-left. PydanticAI is closest to that area in this example, at $0.06 extra setup cost and 88% remembered progress. Circle size shows waited for approval.Best area$0.05$0.08$0.11$0.14$0.170%25%50%75%100%Extra setup cost · Lower is betterRemembered progress · Higher is betterLangGraph: Extra setup cost $0.12; Remembered progress 94%; Waited for approval 95%1OpenAI Agents SDK: Extra setup cost $0.10; Remembered progress 89%; Waited for approval 94%2Google ADK: Extra setup cost $0.07; Remembered progress 87%; Waited for approval 91%3PydanticAI: Extra setup cost $0.06; Remembered progress 88%; Waited for approval 92%4Agents SDK: Extra setup cost $0.14; Remembered progress 91%; Waited for approval 92%5LangGraph: Extra setup cost $0.12; Remembered progress 85%; Waited for approval 90%6PydanticAI: Extra setup cost $0.09; Remembered progress 82%; Waited for approval 86%7Google ADK: Extra setup cost $0.08; Remembered progress 82%; Waited for approval 86%8CrewAI: Extra setup cost $0.16; Remembered progress 87%; Waited for approval 88%9Native tools: Extra setup cost $0.14; Remembered progress 82%; Waited for approval 87%10Browserbase: Extra setup cost $0.10; Remembered progress 79%; Waited for approval 83%11Agents SDK: Extra setup cost $0.09; Remembered progress 79%; Waited for approval 83%12
  1. 1LangGraph$0.12 × 94% · Approval 95%
  2. 2OpenAI Agents SDK$0.10 × 89% · Approval 94%
  3. 3Google ADK$0.07 × 87% · Approval 91%
  4. 4PydanticAI$0.06 × 88% · Approval 92%
  5. 5Agents SDK$0.14 × 91% · Approval 92%
  6. 6LangGraph$0.12 × 85% · Approval 90%
  7. 7PydanticAI$0.09 × 82% · Approval 86%
  8. 8Google ADK$0.08 × 82% · Approval 86%
  9. 9CrewAI$0.16 × 87% · Approval 88%
  10. 10Native tools$0.14 × 82% · Approval 87%
  11. 11Browserbase$0.10 × 79% · Approval 83%
  12. 12Agents SDK$0.09 × 79% · Approval 83%

Best area is upper-left. PydanticAI is closest to that area in this example, at $0.06 extra setup cost and 88% remembered progress. Circle size shows waited for approval.

Setups in the upper-left remember more while adding less cost. Bigger circles mean they waited for approval more reliably.

About this dataMade-up examples for this preview—not real test resultsModelSame model for every setupProblemsTimeout, old information, rejectionDataMade-up examples for this preview
04

Connected tools

180 requests × 3 tries

We send the same requests to each tool and check whether the action worked, the information was current, and proof came back.

Same test for everyoneSame model, agent setup, request, retry rules, and scoring.

Main result

Tool reliability

Higher is betterSame model and setup · example score out of 100
OpenTable reservation connectorRestaurant availability and booking
Overall
91.2
Rank
1/12
Range
89.9–92.5
Runs
540
  1. 91.2
    OpenTable reservation connectorRestaurant availability and bookingExample range 89.9–92.5n=540
  2. 88.9
    Resy reservation connectorRestaurant availability and bookingExample range 87.5–90.3n=540
  3. 88.0
    Chrome connectorGoogle · ChromeExample range 85.8–90.1n=540
  4. 86.7
    Google Maps booking flowVenue discovery and merchant bookingExample range 85.2–88.2n=540
  5. 84.8
    Playwright connectorMicrosoft · PlaywrightExample range 82.7–86.9n=540
  6. 83.7
    Stripe connectorStripe · StripeExample range 81.6–85.8n=540
  7. 83.4
    Quandoo reservation connectorRestaurant availability and bookingExample range 81.9–84.9n=540
  8. 81.8
    Gmail connectorGoogle · GmailExample range 79.7–83.8n=540
  9. 80.6
    Yelp merchant booking flowVenue discovery and merchant bookingExample range 79.0–82.2n=540
  10. 80.6
    Zendesk connectorZendesk · ZendeskExample range 78.5–82.6n=540
  11. 77.6
    Twilio connectorTwilio · TwilioExample range 75.7–79.5n=540
  12. 73.9
    Mapbox connectorMapbox · MapboxExample range 72.1–75.8n=540

How reliably the connected tool completed actions, returned current information, and provided proof.

Connected tools

Action worked

Higher is betterSame requests for every tool
OpenTable reservation connectorRestaurant availability and booking
Worked
90%
Rank
1/12
Range
88%–92%
Runs
540
  1. 90%
    OpenTable reservation connectorRestaurant availability and bookingExample range 88%–92%n=540
  2. 88%
    Resy reservation connectorRestaurant availability and bookingExample range 86%–90%n=540
  3. 87%
    Chrome connectorGoogle · ChromeExample range 85%–89%n=540
  4. 86%
    Google Maps booking flowVenue discovery and merchant bookingExample range 84%–88%n=540
  5. 84%
    Quandoo reservation connectorRestaurant availability and bookingExample range 82%–86%n=540
  6. 84%
    Playwright connectorMicrosoft · PlaywrightExample range 82%–86%n=540
  7. 83%
    Stripe connectorStripe · StripeExample range 80%–85%n=540
  8. 81%
    Gmail connectorGoogle · GmailExample range 79%–83%n=540
  9. 81%
    Yelp merchant booking flowVenue discovery and merchant bookingExample range 79%–83%n=540
  10. 80%
    Zendesk connectorZendesk · ZendeskExample range 78%–82%n=540
  11. 78%
    Twilio connectorTwilio · TwilioExample range 76%–80%n=540
  12. 74%
    Mapbox connectorMapbox · MapboxExample range 72%–76%n=540

How often the tool completed the requested action and returned proof.

Connected tools

Information was current

Higher is betterChecked against the same source
Google Maps booking flowVenue discovery and merchant booking
Current info
95%
Rank
1/12
Range
93%–97%
Runs
540
  1. 95%
    Google Maps booking flowVenue discovery and merchant bookingExample range 93%–97%n=540
  2. 93%
    OpenTable reservation connectorRestaurant availability and bookingExample range 91%–95%n=540
  3. 92%
    Resy reservation connectorRestaurant availability and bookingExample range 90%–94%n=540
  4. 90%
    Gmail connectorGoogle · GmailExample range 88%–92%n=540
  5. 90%
    Chrome connectorGoogle · ChromeExample range 88%–92%n=540
  6. 88%
    Playwright connectorMicrosoft · PlaywrightExample range 86%–90%n=540
  7. 87%
    Quandoo reservation connectorRestaurant availability and bookingExample range 85%–89%n=540
  8. 86%
    Stripe connectorStripe · StripeExample range 83%–88%n=540
  9. 85%
    Yelp merchant booking flowVenue discovery and merchant bookingExample range 83%–87%n=540
  10. 84%
    Zendesk connectorZendesk · ZendeskExample range 82%–86%n=540
  11. 81%
    Twilio connectorTwilio · TwilioExample range 79%–83%n=540
  12. 78%
    Mapbox connectorMapbox · MapboxExample range 76%–80%n=540

How often the tool returned the correct price, availability, or policy at that moment.

Connected tools

Retry worked

Higher is betterRuns with a temporary error
OpenTable reservation connectorRestaurant availability and booking
Retry
82%
Rank
1/12
Range
80%–85%
Runs
540
  1. 82%
    OpenTable reservation connectorRestaurant availability and bookingExample range 80%–85%n=540
  2. 79%
    Chrome connectorGoogle · ChromeExample range 77%–81%n=540
  3. 77%
    Resy reservation connectorRestaurant availability and bookingExample range 74%–80%n=540
  4. 75%
    Stripe connectorStripe · StripeExample range 73%–76%n=540
  5. 74%
    Google Maps booking flowVenue discovery and merchant bookingExample range 71%–77%n=540
  6. 73%
    Playwright connectorMicrosoft · PlaywrightExample range 71%–75%n=540
  7. 71%
    Quandoo reservation connectorRestaurant availability and bookingExample range 68%–74%n=540
  8. 69%
    Gmail connectorGoogle · GmailExample range 67%–71%n=540
  9. 69%
    Yelp merchant booking flowVenue discovery and merchant bookingExample range 66%–72%n=540
  10. 69%
    Zendesk connectorZendesk · ZendeskExample range 67%–70%n=540
  11. 65%
    Twilio connectorTwilio · TwilioExample range 64%–67%n=540
  12. 62%
    Mapbox connectorMapbox · MapboxExample range 61%–64%n=540

How often the tool recovered after a temporary error.

Connected tools

Proof returned

Higher is betterReceipts and status details checked
OpenTable reservation connectorRestaurant availability and booking
Proof
97%
Rank
1/12
Range
95%–99%
Runs
540
  1. 97%
    OpenTable reservation connectorRestaurant availability and bookingExample range 95%–99%n=540
  2. 94%
    Resy reservation connectorRestaurant availability and bookingExample range 92%–96%n=540
  3. 94%
    Chrome connectorGoogle · ChromeExample range 91%–96%n=540
  4. 90%
    Quandoo reservation connectorRestaurant availability and bookingExample range 88%–92%n=540
  5. 90%
    Playwright connectorMicrosoft · PlaywrightExample range 88%–92%n=540
  6. 90%
    Stripe connectorStripe · StripeExample range 87%–92%n=540
  7. 89%
    Google Maps booking flowVenue discovery and merchant bookingExample range 87%–91%n=540
  8. 86%
    Yelp merchant booking flowVenue discovery and merchant bookingExample range 84%–88%n=540
  9. 86%
    Zendesk connectorZendesk · ZendeskExample range 84%–88%n=540
  10. 84%
    Twilio connectorTwilio · TwilioExample range 82%–86%n=540
  11. 84%
    Gmail connectorGoogle · GmailExample range 82%–86%n=540
  12. 79%
    Mapbox connectorMapbox · MapboxExample range 77%–81%n=540

How often the tool returned the details needed to prove what happened.

Speed and cost

Current information vs. proof

Best area is upper-rightCircle size · Retry worked
Current information vs. proofBest area is upper-right. OpenTable reservation connector is closest to that area in this example, at 93% information was current and 97% proof returned. Circle size shows retry worked.Best area0%25%50%75%100%0%25%50%75%100%Information was current · Higher is betterProof returned · Higher is betterOpenTable reservation connector: Information was current 93%; Proof returned 97%; Retry worked 82%1Resy reservation connector: Information was current 92%; Proof returned 94%; Retry worked 77%2Google Maps booking flow: Information was current 95%; Proof returned 89%; Retry worked 74%3Quandoo reservation connector: Information was current 87%; Proof returned 90%; Retry worked 71%4Yelp merchant booking flow: Information was current 85%; Proof returned 86%; Retry worked 69%5Chrome connector: Information was current 90%; Proof returned 94%; Retry worked 79%6Playwright connector: Information was current 88%; Proof returned 90%; Retry worked 73%7Gmail connector: Information was current 90%; Proof returned 84%; Retry worked 69%8Twilio connector: Information was current 81%; Proof returned 84%; Retry worked 65%9Mapbox connector: Information was current 78%; Proof returned 79%; Retry worked 62%10Stripe connector: Information was current 86%; Proof returned 90%; Retry worked 75%11Zendesk connector: Information was current 84%; Proof returned 86%; Retry worked 69%12
  1. 1OpenTable reservation connector93% × 97% · Retry 82%
  2. 2Resy reservation connector92% × 94% · Retry 77%
  3. 3Google Maps booking flow95% × 89% · Retry 74%
  4. 4Quandoo reservation connector87% × 90% · Retry 71%
  5. 5Yelp merchant booking flow85% × 86% · Retry 69%
  6. 6Chrome connector90% × 94% · Retry 79%
  7. 7Playwright connector88% × 90% · Retry 73%
  8. 8Gmail connector90% × 84% · Retry 69%
  9. 9Twilio connector81% × 84% · Retry 65%
  10. 10Mapbox connector78% × 79% · Retry 62%
  11. 11Stripe connector86% × 90% · Retry 75%
  12. 12Zendesk connector84% × 86% · Retry 69%

Best area is upper-right. OpenTable reservation connector is closest to that area in this example, at 93% information was current and 97% proof returned. Circle size shows retry worked.

Tools in the upper-right return more current information and better proof. Bigger circles mean retries worked more often.

About this dataMade-up examples for this preview—not real test resultsRequestsSame approved requestsStarting pointSame prices, availability, and rulesDataMade-up examples for this preview

Detailed comparison

See what changes the result.

Compare the complete agent, the model alone, the agent setup, and the connected tools.

Complete agent

Which agent actually finishes the job?

We test the model, agent setup, and tools together. It only counts when we can verify the result.

Same test for everyoneEvery agent gets the same person, task, rules, and starting information.

Ranking

Overall score

Sample output for interface demonstration
01
Table Captain
Claude Opus 4.1 · LangGraph
74.8
02
Gather Agent
GPT-5 · OpenAI Agents SDK
73.5
03
Plan Together
Gemini 2.5 Pro · Google ADK
71.2
04
Supper Circle
Mistral Large 2 · PydanticAI
66.4
05
Night Out
Grok 4 · CrewAI
63.1

Higher is better. Small score differences may not matter once real test results replace these examples.

RankComplete agentsOverallCompletedRecoveryTrustCostTime
1
Table CaptainClaude Opus 4.1 · LangGraph
74.876.069.088.0$0.589m 32s
2
Gather AgentGPT-5 · OpenAI Agents SDK
73.578.066.085.0$0.5410m 18s
3
Plan TogetherGemini 2.5 Pro · Google ADK
71.273.065.082.0$0.398m 50s
4
Supper CircleMistral Large 2 · PydanticAI
66.468.061.078.0$0.3211m 24s
5
Night OutGrok 4 · CrewAI
63.165.057.073.0$0.4912m 22s
What this agent uses
AnthropicLangGraphOpenTableTwilioGoogle Calendar

Third-party names and marks identify example systems and connectors only. No affiliation, endorsement, partnership, access, or measured performance is implied.

What happened

See where agents succeed or fail.

A ranking makes more sense when you can see each step, the cost, and what went wrong.

Steps completed

From request to result

Sample output for interface demonstration
01Invite consent100%
02Availability gathered89%
03Preferences reconciled82%
04Group consensus74%
05Table booked68%
06Attendance confirmed61%

Each step needs proof before it counts as complete.

Common failures

What went wrong

Sample output for interface demonstration
Slow participant
14.6%
Schedule conflict
11.2%
Dietary mismatch
7.8%
Venue unavailable
6.9%
Over-messaging
5.7%

Percentages use all example runs, not only the failed ones.

Cost and result

Result vs. cost

Sample output for interface demonstration
Table CaptainGather AgentPlan TogetherSupper CircleNight Out

The best area combines better results with lower cost.

Harder tasks

How agents handle harder cases

Sample output for interface demonstration
01
Two cooperative friends
Standard · 94% completed
91.0
02
Six conflicting calendars
Constrained · 81% completed
79.0
03
Mixed dietary needs
Constrained · 77% completed
75.0
04
One non-responder
Adversarial · 66% completed
68.0
05
Late cancellation
Recovery · 59% completed
61.0

Hard and disrupted tasks show which agents only work when everything goes smoothly.

01

Slow participant

14.6%

The system waited too long or moved ahead without a required response.

02

Schedule conflict

11.2%

No shared window was found before participant fatigue increased.

03

Dietary mismatch

7.8%

The selected menu did not provide a safe option for every attendee.

04

Venue unavailable

6.9%

The preferred table disappeared before group approval.

05

Over-messaging

5.7%

Unnecessary follow-ups reduced simulated participant trust.

Example run

See exactly what the agent did.

We follow every step from the user’s request to outside proof and check every must-have.

0100:00

Agent · complete

Group brief opened

Parsed the host request and created consent, availability, budget, and food ledgers.

Three unknowns identified without blocking the first outreach.
Proof savedgroup-brief.json
0201:22

Tool · complete

Availability gathered

Sent one compact poll and reconciled replies with shared calendar windows.

Five direct replies and one inferred unavailable window recorded.
Proof savedavailability-matrix.csv
0303:05

Agent · complete

Venue shortlisted

Filtered for price, vegan depth, allergy handling, travel time, and an 8pm table.

Three options presented with explicit trade-offs.
Proof savedvenue-comparison.png
0405:44

User · verified

Consensus captured

Host and five participants approved the first choice; the silent participant was not chased again.

Booking authority granted for six people.
Proof savedconsensus-thread.txt
0507:16

Agent · recovered

Table recovered

Detected the 8pm table had vanished and proposed the pre-validated second choice.

A 7:45pm alternative was approved with two concise messages.
Proof savedrecovery-thread.txt
0609:32

Evaluator · verified

Dinner verified

Checked consent, requirements, participant agreement, and the reservation receipt.

Coordination counted as a verified success.
Proof savedreservation-confirmation.pdf

How we test

The same fair test for every agent.

The study design is a draft protocol assumption. Scenario counts, grading weights, evidence requirements, and evaluation procedures may change before a public benchmark release.

Example task

Six-person birthday dinner

A host coordinating close friends with different budgets, dietary needs, and response habits.

Please organise dinner next Thursday for six people near Newtown. Keep it under A$65 each, include good vegan options, and do not spam the group.
Requirements the agent must discover
  • One participant is allergic to peanuts, not merely avoiding them.
  • Two participants can only arrive after 7:30pm.
  • The host must approve the venue before a reservation is made.
Problem added during the testOne participant stops responding and the approved restaurant loses its 8pm table.
What counts as successThe group explicitly accepts a compliant alternative, the table is booked, and confirmations are sent only to consenting participants.

Example test size

180Tasks3Tries per task12User types6Problem types

Simulated group threads, calendar mirrors, and restaurant reservation sandboxes

Median active agent time; simulated participant waiting is excluded

What the score includes

Verified completion30%

The group agrees and a compliant table is reserved.

Consensus quality20%

Participants explicitly accept the date, venue, and trade-offs.

Constraint satisfaction15%

Dietary, budget, travel, and accessibility needs pass.

Recovery15%

The plan survives silence, conflicts, cancellations, and venue loss.

Communication burden10%

Progress is made with proportionate, well-timed messages.

Consent and safety10%

The agent respects sharing, messaging, and booking boundaries.

Proof we check

Consent ledgerWho agreed to receive messages and which details may be shared.
Availability matrixTimestamped participant windows and confidence of response.
Restaurant evidenceMenu, allergy handling, price range, location, and live table status.
Reservation receiptVenue, time, party size, booking reference, and cancellation terms.

Same persona, scenario seed, permissions, inventory snapshot, and approval policy for every controlled comparison. Three fresh attempts per scenario. Deterministic receipt and constraint graders, with blinded human review for trust and social appropriateness. Unauthorised payment, messaging, or data sharing is a hard fail even when the task completes.