Agent tests/EVT-05/Preview

Real-world task test · EVT-05

Local Event Discovery

Can an agent find something the user will actually want to do?

We test whether suggestions match the user, feel fresh, fit their time and budget, lead to action, and improve from feedback.

Highlights · EVT-05

Made-up example data
01

Complete agents

1,200 tasks × 3 tries = 3,600 example runs

We test the model, agent setup, and tools together. It only counts when we can verify the result.

Same test for everyoneEvery agent gets the same person, task, rules, and starting information.

Main result

Recommendations acted on

Higher is betterBased on the example runs
Out ThereGPT-5 · OpenAI Agents SDK
Acted on
43%
Rank
1/12
Range
40%–46%
Runs
3,600
  1. 43%
    Out ThereGPT-5 · OpenAI Agents SDKExample range 40%–46%n=3600
  2. 42%
    Weekend SignalGemini 2.5 Pro · Google ADKExample range 40%–44%n=3600
  3. 40%
    Good PlansClaude Opus 4.1 · LangGraphExample range 37%–43%n=3600
  4. 39%
    Compass AgentOpenAI · Agents SDKExample range 37%–40%n=3600
  5. 38%
    Relay ConciergeGoogle · Google ADKExample range 37%–39%n=3600
  6. 36%
    Local ThreadMistral Large 2 · PydanticAIExample range 33%–39%n=3600
  7. 36%
    Scout AIAnthropic · LangGraphExample range 35%–37%n=3600
  8. 35%
    Nova AssistDeepSeek · LangChainExample range 33%–36%n=3600
  9. 34%
    Tonight MaybeGrok 4 · CrewAIExample range 31%–37%n=3600
  10. 32%
    Waypoint ProxAI · Native toolsExample range 30%–33%n=3600
  11. 30%
    Orbit PlannerMistral · PydanticAIExample range 29%–32%n=3600
  12. 27%
    Task PilotMeta · CrewAIExample range 26%–29%n=3600

The user saved, shared, RSVP’d to, or booked the recommended event.

Cost

Cheapest useful recommendation

Lower is betterExample cost in USD
Local ThreadMistral Large 2 · PydanticAI
Cost
$0.17
Rank
1/12
Range
$0.12–$0.22
Runs
3,600
  1. $0.17
    Local ThreadMistral Large 2 · PydanticAIExample range $0.12–$0.22n=3600
  2. $0.21
    Weekend SignalGemini 2.5 Pro · Google ADKExample range $0.18–$0.24n=3600
  3. $0.21
    Orbit PlannerMistral · PydanticAIExample range $0.18–$0.24n=3600
  4. $0.23
    Tonight MaybeGrok 4 · CrewAIExample range $0.18–$0.28n=3600
  5. $0.24
    Compass AgentOpenAI · Agents SDKExample range $0.21–$0.27n=3600
  6. $0.25
    Out ThereGPT-5 · OpenAI Agents SDKExample range $0.21–$0.29n=3600
  7. $0.27
    Good PlansClaude Opus 4.1 · LangGraphExample range $0.23–$0.31n=3600
  8. $0.28
    Nova AssistDeepSeek · LangChainExample range $0.25–$0.31n=3600
  9. $0.30
    Task PilotMeta · CrewAIExample range $0.27–$0.33n=3600
  10. $0.30
    Relay ConciergeGoogle · Google ADKExample range $0.27–$0.33n=3600
  11. $0.32
    Scout AIAnthropic · LangGraphExample range $0.29–$0.35n=3600
  12. $0.38
    Waypoint ProxAI · Native toolsExample range $0.35–$0.41n=3600

The estimated model and tool cost for each completed task.

Speed

Fastest recommendation

Lower is betterLower is better
Weekend SignalGemini 2.5 Pro · Google ADK
Time
1m 36s
Rank
1/12
Range
1m 22s–1m 50s
Runs
3,600
  1. 1m 36s
    Weekend SignalGemini 2.5 Pro · Google ADKExample range 1m 22s–1m 50sn=3600
  2. 1m 43s
    Out ThereGPT-5 · OpenAI Agents SDKExample range 1m 25s–2m 01sn=3600
  3. 1m 48s
    Compass AgentOpenAI · Agents SDKExample range 1m 34s–2m 02sn=3600
  4. 1m 52s
    Good PlansClaude Opus 4.1 · LangGraphExample range 1m 36s–2m 08sn=3600
  5. 2m 04s
    Local ThreadMistral Large 2 · PydanticAIExample range 1m 44s–2m 24sn=3600
  6. 2m 05s
    Relay ConciergeGoogle · Google ADKExample range 1m 51s–2m 19sn=3600
  7. 2m 10s
    Nova AssistDeepSeek · LangChainExample range 1m 56s–2m 24sn=3600
  8. 2m 11s
    Scout AIAnthropic · LangGraphExample range 1m 57s–2m 25sn=3600
  9. 2m 19s
    Tonight MaybeGrok 4 · CrewAIExample range 1m 57s–2m 41sn=3600
  10. 2m 36s
    Orbit PlannerMistral · PydanticAIExample range 2m 22s–2m 50sn=3600
  11. 2m 36s
    Waypoint ProxAI · Native toolsExample range 2m 22s–2m 50sn=3600
  12. 3m 01s
    Task PilotMeta · CrewAIExample range 2m 47s–3m 15sn=3600

How long the agent actively worked on the task.

Complete agents

People who attended

Higher is betterAll recommendations · example range
Weekend SignalGemini 2.5 Pro · Google ADK
Attended
24.0%
Rank
1/12
Range
21.8%–26.2%
Runs
3,600
  1. 24.0%
    Weekend SignalGemini 2.5 Pro · Google ADKExample range 21.8%–26.2%n=3600
  2. 23.8%
    Out ThereGPT-5 · OpenAI Agents SDKExample range 21.6%–26.0%n=3600
  3. 22.9%
    Good PlansClaude Opus 4.1 · LangGraphExample range 20.7%–25.1%n=3600
  4. 20.8%
    Compass AgentOpenAI · Agents SDKExample range 19.4%–22.1%n=3600
  5. 19.1%
    Local ThreadMistral Large 2 · PydanticAIExample range 16.9%–21.3%n=3600
  6. 18.9%
    Relay ConciergeGoogle · Google ADKExample range 17.5%–20.3%n=3600
  7. 18.8%
    Scout AIAnthropic · LangGraphExample range 17.4%–20.2%n=3600
  8. 16.8%
    Tonight MaybeGrok 4 · CrewAIExample range 14.6%–19.0%n=3600
  9. 16.5%
    Nova AssistDeepSeek · LangChainExample range 15.1%–17.9%n=3600
  10. 14.5%
    Waypoint ProxAI · Native toolsExample range 13.1%–15.9%n=3600
  11. 13.3%
    Orbit PlannerMistral · PydanticAIExample range 11.9%–14.7%n=3600
  12. 10.2%
    Task PilotMeta · CrewAIExample range 8.8%–11.6%n=3600

The user booked or RSVP’d, the event fit their calendar, and they attended.

Complete agents

Overall score

Higher is betterExample score out of 100
Weekend SignalGemini 2.5 Pro · Google ADK
Overall
75.9
Rank
1/12
Range
74.2–77.6
Runs
3,600
  1. 75.9
    Weekend SignalGemini 2.5 Pro · Google ADKExample range 74.2–77.6n=3600
  2. 74.7
    Good PlansClaude Opus 4.1 · LangGraphExample range 72.9–76.5n=3600
  3. 73.8
    Out ThereGPT-5 · OpenAI Agents SDKExample range 71.9–75.7n=3600
  4. 72.7
    Compass AgentOpenAI · Agents SDKExample range 70.8–74.5n=3600
  5. 70.6
    Scout AIAnthropic · LangGraphExample range 68.8–72.4n=3600
  6. 68.8
    Relay ConciergeGoogle · Google ADKExample range 67.1–70.6n=3600
  7. 68.6
    Local ThreadMistral Large 2 · PydanticAIExample range 66.7–70.5n=3600
  8. 68.4
    Nova AssistDeepSeek · LangChainExample range 66.7–70.1n=3600
  9. 66.4
    Waypoint ProxAI · Native toolsExample range 64.7–68.0n=3600
  10. 64.9
    Tonight MaybeGrok 4 · CrewAIExample range 62.9–66.9n=3600
  11. 62.8
    Orbit PlannerMistral · PydanticAIExample range 61.2–64.4n=3600
  12. 58.3
    Task PilotMeta · CrewAIExample range 56.8–59.7n=3600

A simple score combining task success, following requirements, recovery, cost, and trust.

Complete agents

Found another good option

Higher is betterRuns where something went wrong
Good PlansClaude Opus 4.1 · LangGraph
Recovery
75%
Rank
1/12
Range
72%–78%
Runs
3,600
  1. 75%
    Good PlansClaude Opus 4.1 · LangGraphExample range 72%–78%n=3600
  2. 72%
    Weekend SignalGemini 2.5 Pro · Google ADKExample range 69%–75%n=3600
  3. 71%
    Scout AIAnthropic · LangGraphExample range 69%–73%n=3600
  4. 69%
    Out ThereGPT-5 · OpenAI Agents SDKExample range 66%–72%n=3600
  5. 69%
    Compass AgentOpenAI · Agents SDKExample range 67%–70%n=3600
  6. 67%
    Waypoint ProxAI · Native toolsExample range 65%–68%n=3600
  7. 65%
    Local ThreadMistral Large 2 · PydanticAIExample range 62%–68%n=3600
  8. 65%
    Nova AssistDeepSeek · LangChainExample range 63%–66%n=3600
  9. 64%
    Relay ConciergeGoogle · Google ADKExample range 62%–66%n=3600
  10. 59%
    Orbit PlannerMistral · PydanticAIExample range 58%–61%n=3600
  11. 59%
    Tonight MaybeGrok 4 · CrewAIExample range 56%–62%n=3600
  12. 52%
    Task PilotMeta · CrewAIExample range 51%–54%n=3600

How often the agent still finished after a price change, tool error, rejection, or timeout.

Complete agents

Followed approval rules

Higher is betterExample score out of 100
Good PlansClaude Opus 4.1 · LangGraph
Trust
91%
Rank
1/12
Range
89%–93%
Runs
3,600
  1. 91%
    Good PlansClaude Opus 4.1 · LangGraphExample range 89%–93%n=3600
  2. 87%
    Weekend SignalGemini 2.5 Pro · Google ADKExample range 85%–89%n=3600
  3. 87%
    Scout AIAnthropic · LangGraphExample range 85%–89%n=3600
  4. 86%
    Out ThereGPT-5 · OpenAI Agents SDKExample range 84%–88%n=3600
  5. 84%
    Compass AgentOpenAI · Agents SDKExample range 82%–86%n=3600
  6. 83%
    Waypoint ProxAI · Native toolsExample range 81%–85%n=3600
  7. 81%
    Relay ConciergeGoogle · Google ADKExample range 79%–83%n=3600
  8. 81%
    Local ThreadMistral Large 2 · PydanticAIExample range 79%–83%n=3600
  9. 80%
    Nova AssistDeepSeek · LangChainExample range 78%–81%n=3600
  10. 75%
    Orbit PlannerMistral · PydanticAIExample range 73%–77%n=3600
  11. 75%
    Tonight MaybeGrok 4 · CrewAIExample range 73%–77%n=3600
  12. 68%
    Task PilotMeta · CrewAIExample range 67%–70%n=3600

Whether the agent asked before important actions, showed its proof, and was honest about uncertainty.

Speed and cost

Time vs. cost

Best area is lower-leftCircle size · Recommendations acted on
Time vs. costBest area is lower-left. Weekend Signal is closest to that area in this example, at 1m 36s fastest recommendation and $0.21 cheapest useful recommendation. Circle size shows recommendations acted on.Best area1m 24s1m 51s2m 19s2m 46s3m 13s$0.14$0.21$0.27$0.34$0.41Fastest recommendation · Lower is betterCheapest useful recommendation · Lower is betterWeekend Signal: Fastest recommendation 1m 36s; Cheapest useful recommendation $0.21; Recommendations acted on 42%1Good Plans: Fastest recommendation 1m 52s; Cheapest useful recommendation $0.27; Recommendations acted on 40%2Out There: Fastest recommendation 1m 43s; Cheapest useful recommendation $0.25; Recommendations acted on 43%3Local Thread: Fastest recommendation 2m 04s; Cheapest useful recommendation $0.17; Recommendations acted on 36%4Tonight Maybe: Fastest recommendation 2m 19s; Cheapest useful recommendation $0.23; Recommendations acted on 34%5Compass Agent: Fastest recommendation 1m 48s; Cheapest useful recommendation $0.24; Recommendations acted on 39%6Scout AI: Fastest recommendation 2m 11s; Cheapest useful recommendation $0.32; Recommendations acted on 36%7Relay Concierge: Fastest recommendation 2m 05s; Cheapest useful recommendation $0.30; Recommendations acted on 38%8Orbit Planner: Fastest recommendation 2m 36s; Cheapest useful recommendation $0.21; Recommendations acted on 30%9Task Pilot: Fastest recommendation 3m 01s; Cheapest useful recommendation $0.30; Recommendations acted on 27%10Nova Assist: Fastest recommendation 2m 10s; Cheapest useful recommendation $0.28; Recommendations acted on 35%11Waypoint Pro: Fastest recommendation 2m 36s; Cheapest useful recommendation $0.38; Recommendations acted on 32%12
  1. 1Weekend Signal1m 36s × $0.21 · Acted on 42%
  2. 2Good Plans1m 52s × $0.27 · Acted on 40%
  3. 3Out There1m 43s × $0.25 · Acted on 43%
  4. 4Local Thread2m 04s × $0.17 · Acted on 36%
  5. 5Tonight Maybe2m 19s × $0.23 · Acted on 34%
  6. 6Compass Agent1m 48s × $0.24 · Acted on 39%
  7. 7Scout AI2m 11s × $0.32 · Acted on 36%
  8. 8Relay Concierge2m 05s × $0.30 · Acted on 38%
  9. 9Orbit Planner2m 36s × $0.21 · Acted on 30%
  10. 10Task Pilot3m 01s × $0.30 · Acted on 27%
  11. 11Nova Assist2m 10s × $0.28 · Acted on 35%
  12. 12Waypoint Pro2m 36s × $0.38 · Acted on 32%

Best area is lower-left. Weekend Signal is closest to that area in this example, at 1m 36s fastest recommendation and $0.21 cheapest useful recommendation. Circle size shows recommendations acted on.

Agents closer to the lower-left finished faster and cost less. Bigger circles led to more actions.

About this dataMade-up examples for this preview—not real test resultsRangeGrouped by userProof requiredReceipt or other outside proofDataMade-up examples for this preview
02

Models only

1200 tasks × 3 tries

Each model sees the same task and options. Tools are turned off, so the model cannot take actions.

Same test for everyoneSame task, options, and scoring. Tools are off.

Main result

Overall reasoning

Higher is betterTools off · example score out of 100
Claude Opus 4.1Anthropic · tools disabled
Overall
86.8
Rank
1/12
Range
85.4–88.2
Runs
3,600
  1. 86.8
    Claude Opus 4.1Anthropic · tools disabledExample range 85.4–88.2n=3600
  2. 85.9
    Gemini 2.5 ProGoogle · tools disabledExample range 84.4–87.4n=3600
  3. 84.7
    GPT-5OpenAI · tools disabledExample range 83.1–86.3n=3600
  4. 83.5
    GPT-4.1OpenAI · tools disabledExample range 81.5–85.6n=3600
  5. 81.8
    Claude 3.7 SonnetAnthropic · tools disabledExample range 79.8–83.8n=3600
  6. 79.8
    Gemini 2.0 FlashGoogle · tools disabledExample range 77.8–81.7n=3600
  7. 79.3
    DeepSeek V3DeepSeek · tools disabledExample range 77.3–81.3n=3600
  8. 77.6
    Grok 3xAI · tools disabledExample range 75.6–79.5n=3600
  9. 77.2
    Mistral Large 2Mistral AI · tools disabledExample range 75.6–78.8n=3600
  10. 74.6
    Grok 4xAI · tools disabledExample range 72.9–76.3n=3600
  11. 71.4
    Mistral LargeMistral · tools disabledExample range 69.6–73.2n=3600
  12. 67.9
    Llama 4 MaverickMeta · tools disabledExample range 66.3–69.6n=3600

How well the model understood the task and made a choice. The model could not use tools or take actions.

Models only

Remembered requirements

Higher is betterSame task information for every model
Claude Opus 4.1Anthropic · tools disabled
Requirements
94%
Rank
1/12
Range
92%–96%
Runs
3,600
  1. 94%
    Claude Opus 4.1Anthropic · tools disabledExample range 92%–96%n=3600
  2. 91%
    Gemini 2.5 ProGoogle · tools disabledExample range 89%–93%n=3600
  3. 91%
    GPT-4.1OpenAI · tools disabledExample range 88%–93%n=3600
  4. 90%
    GPT-5OpenAI · tools disabledExample range 88%–92%n=3600
  5. 87%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 85%–89%n=3600
  6. 87%
    DeepSeek V3DeepSeek · tools disabledExample range 84%–89%n=3600
  7. 85%
    Gemini 2.0 FlashGoogle · tools disabledExample range 83%–87%n=3600
  8. 84%
    Mistral Large 2Mistral AI · tools disabledExample range 82%–86%n=3600
  9. 83%
    Grok 3xAI · tools disabledExample range 81%–85%n=3600
  10. 81%
    Grok 4xAI · tools disabledExample range 79%–83%n=3600
  11. 78%
    Mistral LargeMistral · tools disabledExample range 76%–80%n=3600
  12. 74%
    Llama 4 MaverickMeta · tools disabledExample range 72%–76%n=3600

How often the model remembered the user’s must-haves when making a choice.

Models only

Choice quality

Higher is betterSame options and scoring
Gemini 2.5 ProGoogle · tools disabled
Choice
90%
Rank
1/12
Range
88%–92%
Runs
3,600
  1. 90%
    Gemini 2.5 ProGoogle · tools disabledExample range 88%–92%n=3600
  2. 88%
    Claude Opus 4.1Anthropic · tools disabledExample range 86%–90%n=3600
  3. 88%
    GPT-5OpenAI · tools disabledExample range 86%–90%n=3600
  4. 86%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 84%–88%n=3600
  5. 85%
    GPT-4.1OpenAI · tools disabledExample range 83%–87%n=3600
  6. 83%
    Gemini 2.0 FlashGoogle · tools disabledExample range 81%–85%n=3600
  7. 82%
    Grok 3xAI · tools disabledExample range 80%–84%n=3600
  8. 81%
    DeepSeek V3DeepSeek · tools disabledExample range 78%–83%n=3600
  9. 79%
    Mistral Large 2Mistral AI · tools disabledExample range 77%–81%n=3600
  10. 77%
    Grok 4xAI · tools disabledExample range 75%–79%n=3600
  11. 73%
    Mistral LargeMistral · tools disabledExample range 71%–75%n=3600
  12. 70%
    Llama 4 MaverickMeta · tools disabledExample range 69%–72%n=3600

How good the model’s choice was when every model saw the same options.

Models only

Action plan works

Higher is betterPlans reviewed without running them
GPT-5OpenAI · tools disabled
Plan
86%
Rank
1/12
Range
84%–88%
Runs
3,600
  1. 86%
    GPT-5OpenAI · tools disabledExample range 84%–88%n=3600
  2. 85%
    Gemini 2.5 ProGoogle · tools disabledExample range 83%–87%n=3600
  3. 84%
    Claude Opus 4.1Anthropic · tools disabledExample range 82%–86%n=3600
  4. 81%
    Gemini 2.0 FlashGoogle · tools disabledExample range 79%–83%n=3600
  5. 81%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 79%–83%n=3600
  6. 81%
    GPT-4.1OpenAI · tools disabledExample range 79%–83%n=3600
  7. 77%
    Grok 3xAI · tools disabledExample range 75%–79%n=3600
  8. 77%
    DeepSeek V3DeepSeek · tools disabledExample range 75%–78%n=3600
  9. 75%
    Mistral Large 2Mistral AI · tools disabledExample range 73%–78%n=3600
  10. 73%
    Grok 4xAI · tools disabledExample range 70%–76%n=3600
  11. 69%
    Mistral LargeMistral · tools disabledExample range 67%–71%n=3600
  12. 66%
    Llama 4 MaverickMeta · tools disabledExample range 65%–68%n=3600

How often the proposed steps were possible, in the right order, and safe to follow.

Models only

Honest about uncertainty

Higher is betterSame scoring for every model
Claude Opus 4.1Anthropic · tools disabled
Honesty
92%
Rank
1/12
Range
90%–94%
Runs
3,600
  1. 92%
    Claude Opus 4.1Anthropic · tools disabledExample range 90%–94%n=3600
  2. 89%
    GPT-4.1OpenAI · tools disabledExample range 87%–91%n=3600
  3. 87%
    Gemini 2.5 ProGoogle · tools disabledExample range 85%–89%n=3600
  4. 86%
    GPT-5OpenAI · tools disabledExample range 84%–88%n=3600
  5. 85%
    DeepSeek V3DeepSeek · tools disabledExample range 82%–87%n=3600
  6. 83%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 81%–85%n=3600
  7. 81%
    Gemini 2.0 FlashGoogle · tools disabledExample range 79%–83%n=3600
  8. 81%
    Mistral Large 2Mistral AI · tools disabledExample range 79%–83%n=3600
  9. 79%
    Grok 3xAI · tools disabledExample range 77%–81%n=3600
  10. 75%
    Mistral LargeMistral · tools disabledExample range 73%–77%n=3600
  11. 74%
    Grok 4xAI · tools disabledExample range 72%–76%n=3600
  12. 67%
    Llama 4 MaverickMeta · tools disabledExample range 66%–69%n=3600

Whether the model used the available proof and avoided claiming success when it was unsure.

Speed and cost

Good choice vs. workable plan

Best area is upper-rightCircle size · Honest about uncertainty
Good choice vs. workable planBest area is upper-right. Gemini 2.5 Pro is closest to that area in this example, at 85% action plan works and 90% choice quality. Circle size shows honest about uncertainty.Best area0%25%50%75%100%0%25%50%75%100%Action plan works · Higher is betterChoice quality · Higher is betterClaude Opus 4.1: Action plan works 84%; Choice quality 88%; Honest about uncertainty 92%1Gemini 2.5 Pro: Action plan works 85%; Choice quality 90%; Honest about uncertainty 87%2GPT-5: Action plan works 86%; Choice quality 88%; Honest about uncertainty 86%3Mistral Large 2: Action plan works 75%; Choice quality 79%; Honest about uncertainty 81%4Grok 4: Action plan works 73%; Choice quality 77%; Honest about uncertainty 74%5GPT-4.1: Action plan works 81%; Choice quality 85%; Honest about uncertainty 89%6Claude 3.7 Sonnet: Action plan works 81%; Choice quality 86%; Honest about uncertainty 83%7Gemini 2.0 Flash: Action plan works 81%; Choice quality 83%; Honest about uncertainty 81%8Mistral Large: Action plan works 69%; Choice quality 73%; Honest about uncertainty 75%9Llama 4 Maverick: Action plan works 66%; Choice quality 70%; Honest about uncertainty 67%10DeepSeek V3: Action plan works 77%; Choice quality 81%; Honest about uncertainty 85%11Grok 3: Action plan works 77%; Choice quality 82%; Honest about uncertainty 79%12
  1. 1Claude Opus 4.184% × 88% · Honesty 92%
  2. 2Gemini 2.5 Pro85% × 90% · Honesty 87%
  3. 3GPT-586% × 88% · Honesty 86%
  4. 4Mistral Large 275% × 79% · Honesty 81%
  5. 5Grok 473% × 77% · Honesty 74%
  6. 6GPT-4.181% × 85% · Honesty 89%
  7. 7Claude 3.7 Sonnet81% × 86% · Honesty 83%
  8. 8Gemini 2.0 Flash81% × 83% · Honesty 81%
  9. 9Mistral Large69% × 73% · Honesty 75%
  10. 10Llama 4 Maverick66% × 70% · Honesty 67%
  11. 11DeepSeek V377% × 81% · Honesty 85%
  12. 12Grok 377% × 82% · Honesty 79%

Best area is upper-right. Gemini 2.5 Pro is closest to that area in this example, at 85% action plan works and 90% choice quality. Circle size shows honest about uncertainty.

Models in the upper-right make better choices and produce plans that are more likely to work. Bigger circles mean the model was more honest about uncertainty.

About this dataMade-up examples for this preview—not real test resultsActionsTurned offOptionsSame options for every modelDataMade-up examples for this preview
03

Agent setup

1200 tasks × 5 problem types

We keep the model and tools the same and change only the workflow that runs them.

Same test for everyoneSame model, tools, task, and problems.

Main result

Agent setup score

Higher is betterSame model and tools · example score out of 100
LangGraphLangChain · frozen model and adapters
Overall
86.4
Rank
1/12
Range
84.9–87.9
Runs
3,600
  1. 86.4
    LangGraphLangChain · frozen model and adaptersExample range 84.9–87.9n=3600
  2. 84.8
    Google ADKGoogle · frozen model and adaptersExample range 83.2–86.4n=3600
  3. 83.9
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 82.2–85.6n=3600
  4. 83.2
    Agents SDKOpenAI · Agents SDKExample range 81.1–85.2n=3600
  5. 80.7
    LangGraphLangChain · LangGraphExample range 78.7–82.7n=3600
  6. 79.8
    CrewAICrewAI · CrewAIExample range 77.8–81.7n=3600
  7. 79.6
    PydanticAIPydantic · frozen model and adaptersExample range 77.9–81.3n=3600
  8. 79.0
    PydanticAIPydantic · PydanticAIExample range 77.0–80.9n=3600
  9. 77.3
    Native toolsAnthropic · Native toolsExample range 75.4–79.2n=3600
  10. 75.6
    BrowserbaseBrowserbase · BrowserbaseExample range 73.7–77.4n=3600
  11. 73.8
    Google ADKGoogle · Google ADKExample range 72.0–75.6n=3600
  12. 70.4
    Agents SDKOpenAI · Agents SDKExample range 68.6–72.2n=3600

How well the workflow ran the task when the model and tools stayed the same.

Agent setup

Ran the plan correctly

Higher is betterSame approved plan
Google ADKGoogle · frozen model and adapters
Execution
85%
Rank
1/12
Range
83%–87%
Runs
3,600
  1. 85%
    Google ADKGoogle · frozen model and adaptersExample range 83%–87%n=3600
  2. 84%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 82%–86%n=3600
  3. 82%
    LangGraphLangChain · frozen model and adaptersExample range 80%–84%n=3600
  4. 81%
    LangGraphLangChain · LangGraphExample range 79%–83%n=3600
  5. 79%
    PydanticAIPydantic · PydanticAIExample range 77%–81%n=3600
  6. 79%
    Agents SDKOpenAI · Agents SDKExample range 77%–81%n=3600
  7. 78%
    Native toolsAnthropic · Native toolsExample range 76%–79%n=3600
  8. 77%
    PydanticAIPydantic · frozen model and adaptersExample range 75%–79%n=3600
  9. 76%
    BrowserbaseBrowserbase · BrowserbaseExample range 74%–78%n=3600
  10. 75%
    CrewAICrewAI · CrewAIExample range 73%–77%n=3600
  11. 71%
    Google ADKGoogle · Google ADKExample range 69%–73%n=3600
  12. 68%
    Agents SDKOpenAI · Agents SDKExample range 66%–69%n=3600

How often the setup followed the approved steps in the right order without losing progress.

Agent setup

Recovered from an error

Higher is betterRuns where something went wrong
LangGraphLangChain · frozen model and adapters
Recovery
81%
Rank
1/12
Range
78%–84%
Runs
3,600
  1. 81%
    LangGraphLangChain · frozen model and adaptersExample range 78%–84%n=3600
  2. 78%
    Agents SDKOpenAI · Agents SDKExample range 76%–80%n=3600
  3. 77%
    Google ADKGoogle · frozen model and adaptersExample range 74%–80%n=3600
  4. 75%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 72%–78%n=3600
  5. 74%
    CrewAICrewAI · CrewAIExample range 72%–76%n=3600
  6. 74%
    PydanticAIPydantic · frozen model and adaptersExample range 71%–77%n=3600
  7. 73%
    LangGraphLangChain · LangGraphExample range 71%–75%n=3600
  8. 70%
    PydanticAIPydantic · PydanticAIExample range 68%–72%n=3600
  9. 70%
    Native toolsAnthropic · Native toolsExample range 68%–71%n=3600
  10. 68%
    Google ADKGoogle · Google ADKExample range 66%–70%n=3600
  11. 67%
    BrowserbaseBrowserbase · BrowserbaseExample range 65%–68%n=3600
  12. 65%
    Agents SDKOpenAI · Agents SDKExample range 63%–66%n=3600

How often the setup recovered from a timeout, outdated information, price change, or rejection.

Agent setup

Remembered progress

Higher is betterProgress record checked
LangGraphLangChain · frozen model and adapters
Memory
96%
Rank
1/12
Range
94%–98%
Runs
3,600
  1. 96%
    LangGraphLangChain · frozen model and adaptersExample range 94%–98%n=3600
  2. 93%
    Google ADKGoogle · frozen model and adaptersExample range 91%–95%n=3600
  3. 93%
    Agents SDKOpenAI · Agents SDKExample range 90%–95%n=3600
  4. 92%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 90%–94%n=3600
  5. 91%
    PydanticAIPydantic · frozen model and adaptersExample range 89%–93%n=3600
  6. 89%
    CrewAICrewAI · CrewAIExample range 87%–92%n=3600
  7. 89%
    LangGraphLangChain · LangGraphExample range 87%–91%n=3600
  8. 87%
    PydanticAIPydantic · PydanticAIExample range 85%–89%n=3600
  9. 86%
    Native toolsAnthropic · Native toolsExample range 83%–88%n=3600
  10. 85%
    Google ADKGoogle · Google ADKExample range 83%–87%n=3600
  11. 84%
    BrowserbaseBrowserbase · BrowserbaseExample range 82%–86%n=3600
  12. 82%
    Agents SDKOpenAI · Agents SDKExample range 80%–84%n=3600

Whether the setup remembered requirements, approvals, evidence, and completed steps.

Agent setup

Waited for approval

Higher is betterApproval checks
LangGraphLangChain · frozen model and adapters
Approval
95%
Rank
1/12
Range
93%–97%
Runs
3,600
  1. 95%
    LangGraphLangChain · frozen model and adaptersExample range 93%–97%n=3600
  2. 94%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 92%–96%n=3600
  3. 93%
    Google ADKGoogle · frozen model and adaptersExample range 91%–95%n=3600
  4. 92%
    PydanticAIPydantic · frozen model and adaptersExample range 90%–94%n=3600
  5. 92%
    Agents SDKOpenAI · Agents SDKExample range 89%–94%n=3600
  6. 89%
    PydanticAIPydantic · PydanticAIExample range 87%–91%n=3600
  7. 89%
    LangGraphLangChain · LangGraphExample range 87%–91%n=3600
  8. 88%
    CrewAICrewAI · CrewAIExample range 86%–91%n=3600
  9. 86%
    Google ADKGoogle · Google ADKExample range 84%–88%n=3600
  10. 86%
    BrowserbaseBrowserbase · BrowserbaseExample range 84%–88%n=3600
  11. 86%
    Native toolsAnthropic · Native toolsExample range 83%–88%n=3600
  12. 83%
    Agents SDKOpenAI · Agents SDKExample range 81%–85%n=3600

How often the setup asked the user before an important or hard-to-reverse action.

Cost

Extra setup cost

Lower is betterExample USD above the baseline
PydanticAIPydantic · frozen model and adapters
Extra cost
$0.03
Rank
1/12
Range
$0.00–$0.06
Runs
3,600
  1. $0.03
    PydanticAIPydantic · frozen model and adaptersExample range $0.00–$0.06n=3600
  2. $0.04
    Google ADKGoogle · Google ADKExample range $0.01–$0.07n=3600
  3. $0.04
    Google ADKGoogle · frozen model and adaptersExample range $0.02–$0.06n=3600
  4. $0.04
    Agents SDKOpenAI · Agents SDKExample range $0.01–$0.07n=3600
  5. $0.05
    LangGraphLangChain · LangGraphExample range $0.02–$0.08n=3600
  6. $0.05
    LangGraphLangChain · frozen model and adaptersExample range $0.03–$0.07n=3600
  7. $0.05
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range $0.02–$0.08n=3600
  8. $0.05
    Native toolsAnthropic · Native toolsExample range $0.02–$0.08n=3600
  9. $0.06
    Agents SDKOpenAI · Agents SDKExample range $0.03–$0.09n=3600
  10. $0.06
    PydanticAIPydantic · PydanticAIExample range $0.03–$0.09n=3600
  11. $0.07
    CrewAICrewAI · CrewAIExample range $0.04–$0.10n=3600
  12. $0.07
    BrowserbaseBrowserbase · BrowserbaseExample range $0.04–$0.10n=3600

The extra cost added by the agent setup for each completed task.

Speed and cost

Memory vs. extra cost

Best area is upper-leftCircle size · Waited for approval
Memory vs. extra costBest area is upper-left. PydanticAI is closest to that area in this example, at $0.03 extra setup cost and 91% remembered progress. Circle size shows waited for approval.Best area$0.02$0.03$0.05$0.07$0.080%25%50%75%100%Extra setup cost · Lower is betterRemembered progress · Higher is betterLangGraph: Extra setup cost $0.05; Remembered progress 96%; Waited for approval 95%1Google ADK: Extra setup cost $0.04; Remembered progress 93%; Waited for approval 93%2OpenAI Agents SDK: Extra setup cost $0.05; Remembered progress 92%; Waited for approval 94%3PydanticAI: Extra setup cost $0.03; Remembered progress 91%; Waited for approval 92%4Agents SDK: Extra setup cost $0.06; Remembered progress 93%; Waited for approval 92%5LangGraph: Extra setup cost $0.05; Remembered progress 89%; Waited for approval 89%6PydanticAI: Extra setup cost $0.06; Remembered progress 87%; Waited for approval 89%7Google ADK: Extra setup cost $0.04; Remembered progress 85%; Waited for approval 86%8CrewAI: Extra setup cost $0.07; Remembered progress 89%; Waited for approval 88%9Native tools: Extra setup cost $0.05; Remembered progress 86%; Waited for approval 86%10Browserbase: Extra setup cost $0.07; Remembered progress 84%; Waited for approval 86%11Agents SDK: Extra setup cost $0.04; Remembered progress 82%; Waited for approval 83%12
  1. 1LangGraph$0.05 × 96% · Approval 95%
  2. 2Google ADK$0.04 × 93% · Approval 93%
  3. 3OpenAI Agents SDK$0.05 × 92% · Approval 94%
  4. 4PydanticAI$0.03 × 91% · Approval 92%
  5. 5Agents SDK$0.06 × 93% · Approval 92%
  6. 6LangGraph$0.05 × 89% · Approval 89%
  7. 7PydanticAI$0.06 × 87% · Approval 89%
  8. 8Google ADK$0.04 × 85% · Approval 86%
  9. 9CrewAI$0.07 × 89% · Approval 88%
  10. 10Native tools$0.05 × 86% · Approval 86%
  11. 11Browserbase$0.07 × 84% · Approval 86%
  12. 12Agents SDK$0.04 × 82% · Approval 83%

Best area is upper-left. PydanticAI is closest to that area in this example, at $0.03 extra setup cost and 91% remembered progress. Circle size shows waited for approval.

Setups in the upper-left remember more while adding less cost. Bigger circles mean they waited for approval more reliably.

About this dataMade-up examples for this preview—not real test resultsModelSame model for every setupProblemsTimeout, old information, rejectionDataMade-up examples for this preview
04

Connected tools

1200 requests × 3 tries

We send the same requests to each tool and check whether the action worked, the information was current, and proof came back.

Same test for everyoneSame model, agent setup, request, retry rules, and scoring.

Main result

Tool reliability

Higher is betterSame model and setup · example score out of 100
Eventbrite discovery connectorCatalog search, availability, and ticket handoff
Overall
91.1
Rank
1/12
Range
89.8–92.4
Runs
3,600
  1. 91.1
    Eventbrite discovery connectorCatalog search, availability, and ticket handoffExample range 89.8–92.4n=3600
  2. 89.4
    Meetup discovery connectorCommunity event search and RSVPExample range 88.0–90.8n=3600
  3. 87.8
    Chrome connectorGoogle · ChromeExample range 85.7–90.0n=3600
  4. 87.2
    DICE event connectorMusic discovery, availability, and ticket handoffExample range 85.7–88.7n=3600
  5. 85.3
    Playwright connectorMicrosoft · PlaywrightExample range 83.2–87.4n=3600
  6. 84.9
    Humanitix event connectorLocal event search and ticket handoffExample range 83.4–86.4n=3600
  7. 83.6
    Stripe connectorStripe · StripeExample range 81.5–85.7n=3600
  8. 82.8
    Fever discovery connectorCurated experience search and booking handoffExample range 81.2–84.4n=3600
  9. 82.3
    Gmail connectorGoogle · GmailExample range 80.2–84.3n=3600
  10. 81.1
    Zendesk connectorZendesk · ZendeskExample range 79.0–83.1n=3600
  11. 79.1
    Twilio connectorTwilio · TwilioExample range 77.1–81.1n=3600
  12. 76.1
    Mapbox connectorMapbox · MapboxExample range 74.2–78.1n=3600

How reliably the connected tool completed actions, returned current information, and provided proof.

Connected tools

Action worked

Higher is betterSame requests for every tool
Eventbrite discovery connectorCatalog search, availability, and ticket handoff
Worked
91%
Rank
1/12
Range
89%–93%
Runs
3,600
  1. 91%
    Eventbrite discovery connectorCatalog search, availability, and ticket handoffExample range 89%–93%n=3600
  2. 89%
    Meetup discovery connectorCommunity event search and RSVPExample range 87%–91%n=3600
  3. 88%
    DICE event connectorMusic discovery, availability, and ticket handoffExample range 86%–90%n=3600
  4. 88%
    Chrome connectorGoogle · ChromeExample range 86%–90%n=3600
  5. 85%
    Humanitix event connectorLocal event search and ticket handoffExample range 83%–87%n=3600
  6. 85%
    Playwright connectorMicrosoft · PlaywrightExample range 83%–87%n=3600
  7. 84%
    Stripe connectorStripe · StripeExample range 81%–86%n=3600
  8. 83%
    Gmail connectorGoogle · GmailExample range 81%–85%n=3600
  9. 83%
    Fever discovery connectorCurated experience search and booking handoffExample range 81%–85%n=3600
  10. 81%
    Zendesk connectorZendesk · ZendeskExample range 79%–83%n=3600
  11. 79%
    Twilio connectorTwilio · TwilioExample range 77%–81%n=3600
  12. 76%
    Mapbox connectorMapbox · MapboxExample range 74%–78%n=3600

How often the tool completed the requested action and returned proof.

Connected tools

Information was current

Higher is betterChecked against the same source
Eventbrite discovery connectorCatalog search, availability, and ticket handoff
Current info
95%
Rank
1/12
Range
93%–97%
Runs
3,600
  1. 95%
    Eventbrite discovery connectorCatalog search, availability, and ticket handoffExample range 93%–97%n=3600
  2. 93%
    Meetup discovery connectorCommunity event search and RSVPExample range 91%–95%n=3600
  3. 92%
    DICE event connectorMusic discovery, availability, and ticket handoffExample range 90%–94%n=3600
  4. 92%
    Chrome connectorGoogle · ChromeExample range 89%–94%n=3600
  5. 89%
    Humanitix event connectorLocal event search and ticket handoffExample range 87%–91%n=3600
  6. 89%
    Playwright connectorMicrosoft · PlaywrightExample range 87%–91%n=3600
  7. 88%
    Fever discovery connectorCurated experience search and booking handoffExample range 86%–90%n=3600
  8. 88%
    Stripe connectorStripe · StripeExample range 85%–90%n=3600
  9. 87%
    Gmail connectorGoogle · GmailExample range 85%–89%n=3600
  10. 85%
    Zendesk connectorZendesk · ZendeskExample range 83%–87%n=3600
  11. 83%
    Twilio connectorTwilio · TwilioExample range 81%–85%n=3600
  12. 81%
    Mapbox connectorMapbox · MapboxExample range 79%–83%n=3600

How often the tool returned the correct price, availability, or policy at that moment.

Connected tools

Retry worked

Higher is betterRuns with a temporary error
Eventbrite discovery connectorCatalog search, availability, and ticket handoff
Retry
80%
Rank
1/12
Range
78%–83%
Runs
3,600
  1. 80%
    Eventbrite discovery connectorCatalog search, availability, and ticket handoffExample range 78%–83%n=3600
  2. 79%
    Meetup discovery connectorCommunity event search and RSVPExample range 76%–82%n=3600
  3. 77%
    Chrome connectorGoogle · ChromeExample range 75%–79%n=3600
  4. 75%
    Playwright connectorMicrosoft · PlaywrightExample range 73%–77%n=3600
  5. 74%
    DICE event connectorMusic discovery, availability, and ticket handoffExample range 71%–77%n=3600
  6. 73%
    Humanitix event connectorLocal event search and ticket handoffExample range 70%–76%n=3600
  7. 73%
    Stripe connectorStripe · StripeExample range 71%–74%n=3600
  8. 71%
    Zendesk connectorZendesk · ZendeskExample range 69%–72%n=3600
  9. 70%
    Fever discovery connectorCurated experience search and booking handoffExample range 67%–73%n=3600
  10. 69%
    Gmail connectorGoogle · GmailExample range 67%–71%n=3600
  11. 67%
    Twilio connectorTwilio · TwilioExample range 66%–69%n=3600
  12. 63%
    Mapbox connectorMapbox · MapboxExample range 62%–65%n=3600

How often the tool recovered after a temporary error.

Connected tools

Proof returned

Higher is betterReceipts and status details checked
Eventbrite discovery connectorCatalog search, availability, and ticket handoff
Proof
96%
Rank
1/12
Range
94%–98%
Runs
3,600
  1. 96%
    Eventbrite discovery connectorCatalog search, availability, and ticket handoffExample range 94%–98%n=3600
  2. 94%
    Meetup discovery connectorCommunity event search and RSVPExample range 92%–96%n=3600
  3. 93%
    Chrome connectorGoogle · ChromeExample range 90%–95%n=3600
  4. 91%
    DICE event connectorMusic discovery, availability, and ticket handoffExample range 89%–93%n=3600
  5. 90%
    Humanitix event connectorLocal event search and ticket handoffExample range 88%–92%n=3600
  6. 90%
    Playwright connectorMicrosoft · PlaywrightExample range 88%–92%n=3600
  7. 89%
    Stripe connectorStripe · StripeExample range 86%–91%n=3600
  8. 88%
    Fever discovery connectorCurated experience search and booking handoffExample range 86%–90%n=3600
  9. 86%
    Gmail connectorGoogle · GmailExample range 84%–88%n=3600
  10. 86%
    Zendesk connectorZendesk · ZendeskExample range 84%–88%n=3600
  11. 84%
    Twilio connectorTwilio · TwilioExample range 82%–86%n=3600
  12. 81%
    Mapbox connectorMapbox · MapboxExample range 79%–83%n=3600

How often the tool returned the details needed to prove what happened.

Speed and cost

Current information vs. proof

Best area is upper-rightCircle size · Retry worked
Current information vs. proofBest area is upper-right. Eventbrite discovery connector is closest to that area in this example, at 95% information was current and 96% proof returned. Circle size shows retry worked.Best area0%25%50%75%100%0%25%50%75%100%Information was current · Higher is betterProof returned · Higher is betterEventbrite discovery connector: Information was current 95%; Proof returned 96%; Retry worked 80%1Meetup discovery connector: Information was current 93%; Proof returned 94%; Retry worked 79%2DICE event connector: Information was current 92%; Proof returned 91%; Retry worked 74%3Humanitix event connector: Information was current 89%; Proof returned 90%; Retry worked 73%4Fever discovery connector: Information was current 88%; Proof returned 88%; Retry worked 70%5Chrome connector: Information was current 92%; Proof returned 93%; Retry worked 77%6Playwright connector: Information was current 89%; Proof returned 90%; Retry worked 75%7Gmail connector: Information was current 87%; Proof returned 86%; Retry worked 69%8Twilio connector: Information was current 83%; Proof returned 84%; Retry worked 67%9Mapbox connector: Information was current 81%; Proof returned 81%; Retry worked 63%10Stripe connector: Information was current 88%; Proof returned 89%; Retry worked 73%11Zendesk connector: Information was current 85%; Proof returned 86%; Retry worked 71%12
  1. 1Eventbrite discovery connector95% × 96% · Retry 80%
  2. 2Meetup discovery connector93% × 94% · Retry 79%
  3. 3DICE event connector92% × 91% · Retry 74%
  4. 4Humanitix event connector89% × 90% · Retry 73%
  5. 5Fever discovery connector88% × 88% · Retry 70%
  6. 6Chrome connector92% × 93% · Retry 77%
  7. 7Playwright connector89% × 90% · Retry 75%
  8. 8Gmail connector87% × 86% · Retry 69%
  9. 9Twilio connector83% × 84% · Retry 67%
  10. 10Mapbox connector81% × 81% · Retry 63%
  11. 11Stripe connector88% × 89% · Retry 73%
  12. 12Zendesk connector85% × 86% · Retry 71%

Best area is upper-right. Eventbrite discovery connector is closest to that area in this example, at 95% information was current and 96% proof returned. Circle size shows retry worked.

Tools in the upper-right return more current information and better proof. Bigger circles mean retries worked more often.

About this dataMade-up examples for this preview—not real test resultsRequestsSame approved requestsStarting pointSame prices, availability, and rulesDataMade-up examples for this preview

Detailed comparison

See what changes the result.

Compare the complete agent, the model alone, the agent setup, and the connected tools.

Complete agent

Which agent actually finishes the job?

We test the model, agent setup, and tools together. It only counts when we can verify the result.

Same test for everyoneEvery agent gets the same person, task, rules, and starting information.

Ranking

Overall score

Sample output for interface demonstration
01
Weekend Signal
Gemini 2.5 Pro · Google ADK
75.9
02
Good Plans
Claude Opus 4.1 · LangGraph
74.7
03
Out There
GPT-5 · OpenAI Agents SDK
73.8
04
Local Thread
Mistral Large 2 · PydanticAI
68.6
05
Tonight Maybe
Grok 4 · CrewAI
64.9

Higher is better. Small score differences may not matter once real test results replace these examples.

RankComplete agentsOverallCompletedRecoveryTrustCostTime
1
Weekend SignalGemini 2.5 Pro · Google ADK
75.942.072.087.0$0.211m 36s
2
Good PlansClaude Opus 4.1 · LangGraph
74.740.075.091.0$0.271m 52s
3
Out ThereGPT-5 · OpenAI Agents SDK
73.843.069.086.0$0.251m 43s
4
Local ThreadMistral Large 2 · PydanticAI
68.636.065.081.0$0.172m 04s
5
Tonight MaybeGrok 4 · CrewAI
64.934.059.075.0$0.232m 19s
What this agent uses
GoogleGoogle ADKEventbriteGoogle MapsGoogle Calendar

Third-party names and marks identify example systems and connectors only. No affiliation, endorsement, partnership, access, or measured performance is implied.

What happened

See where agents succeed or fail.

A ranking makes more sense when you can see each step, the cost, and what went wrong.

Steps completed

From request to result

Sample output for interface demonstration
01Recommendation shown100%
02Details opened67%
03Saved or shared42%
04Ticket booked29%
05Event attended24%
06Positive follow-up16%

Each step needs proof before it counts as complete.

Common failures

What went wrong

Sample output for interface demonstration
Generic match
13.8%
Sold out
9.1%
Calendar conflict
7.7%
Novelty miss
6.8%
Repeated suggestion
4.9%

Percentages use all example runs, not only the failed ones.

Cost and result

Result vs. cost

Sample output for interface demonstration
Weekend SignalGood PlansOut ThereLocal ThreadTonight Maybe

The best area combines better results with lower cost.

Harder tasks

How agents handle harder cases

Sample output for interface demonstration
01
Known-preference match
Standard · 51% completed
89.0
02
Cold-start discovery
Constrained · 43% completed
78.0
03
Novel but relevant
Adversarial · 37% completed
72.0
04
Sold-out recovery
Recovery · 34% completed
67.0
05
Four-week adaptation
Longitudinal · 39% completed
63.0

Hard and disrupted tasks show which agents only work when everything goes smoothly.

01

Generic match

13.8%

The suggestion matched a broad category but not the person’s actual taste.

02

Sold out

9.1%

The recommendation was relevant but no longer actionable.

03

Calendar conflict

7.7%

The event overlapped a known commitment or travel time.

04

Novelty miss

6.8%

Exploration drifted beyond the user’s tolerance without explanation.

05

Repeated suggestion

4.9%

The system resurfaced a declined event or ignored prior feedback.

Example run

See exactly what the agent did.

We follow every step from the user’s request to outside proof and check every must-have.

0100:00

Agent · complete

Taste model loaded

Combined the request with prior accepts, rejects, budget, access, and social comfort.

Three stable preferences and two exploration bounds recorded.
Proof savedpreference-ledger.json
0200:17

Tool · complete

Catalog searched

Queried local events and checked transit, calendar conflicts, price, and live capacity.

Fifty-six events reduced to six eligible candidates.
Proof savedevent-catalog-snapshot.json
0300:43

Agent · complete

Recommendation explained

Presented one strong match and two alternatives with relevance and novelty reasons.

User opened the small gallery performance.
Proof savedrecommendation-card.png
0401:01

Tool · complete

Sell-out detected

Rechecked availability before ticket handoff.

The preferred event had sold out.
Proof savedavailability-update.json
0501:22

Agent · recovered

Alternative recovered

Promoted the pre-validated listening-room event and requested booking approval.

User approved a A$32 ticket and calendar entry.
Proof savedapproval-message.txt
0601:36

Evaluator · verified

Outcome verified

Matched booking, calendar, simulated attendance, rating, and preference update.

Recommendation counted as a positive longitudinal outcome.
Proof savedfollow-up-record.json

How we test

The same fair test for every agent.

The study design is a draft protocol assumption. Scenario counts, grading weights, evidence requirements, and evaluation procedures may change before a public benchmark release.

Example task

A genuinely new Friday plan

A 29-year-old new to Sydney who likes live music and design, avoids crowded clubs, has a A$45 budget, and wants to meet people without forced networking.

Find me something for Friday night that feels social but not awkward. I want to try something new, stay near a train line, and spend under A$45.
Requirements the agent must discover
  • The persona previously rejected two large standing-room concerts.
  • A 7pm work commitment makes one otherwise strong option impossible.
  • The agent may add nothing to the calendar without explicit approval.
Problem added during the testThe first-choice event sells out after the user opens its details.
What counts as successA novel, compliant alternative is booked in simulation, added with approval, attended, and rated positively in the follow-up.

Example test size

1200Tasks3Tries per task200User types5Problem types

Four simulated weeks of event catalogs, profiles, calendars, and feedback

Median active handling time per recommendation cycle

What the score includes

Downstream action25%

The user saves, shares, books, or attends the recommendation.

Preference fit20%

The event matches stated and learned interests and constraints.

Follow-through20%

The system supports booking, calendar fit, reminders, and recovery.

Calibration and learning15%

Confidence and later recommendations improve with feedback.

Novelty10%

The system expands taste without becoming irrelevant.

Trust and safety10%

The system respects budget, access, privacy, and contact boundaries.

Proof we check

Preference ledgerStated, inferred, accepted, and rejected preferences with provenance.
Event snapshotDescription, time, price, availability, access, and location evidence.
Action recordOpen, save, share, booking, calendar, and simulated attendance events.
Feedback recordPost-event rating, explanation, and resulting preference updates.

Same persona, scenario seed, permissions, inventory snapshot, and approval policy for every controlled comparison. Three fresh attempts per scenario. Deterministic receipt and constraint graders, with blinded human review for trust and social appropriateness. Unauthorised payment, messaging, or data sharing is a hard fail even when the task completes.