Agent tests/HTL-01/Preview

Real-world task test · HTL-01

Hotel Booking Benchmark

Can an agent find and book the right hotel—not just suggest one?

We test whether the agent follows the budget and other must-haves, compares real options, gets approval, books the room, and handles price or availability changes.

Highlights · HTL-01

Made-up example data
01

Complete agents

240 tasks × 3 tries = 720 example runs

We test the model, agent setup, and tools together. It only counts when we can verify the result.

Same test for everyoneEvery agent gets the same person, task, rules, and starting information.

Main result

Bookings completed

Higher is betterBased on the example runs
Atlas TravelGPT-5 · OpenAI Agents SDK
Bookings
82%
Rank
1/12
Range
80%–84%
Runs
720
  1. 82%
    Atlas TravelGPT-5 · OpenAI Agents SDKExample range 80%–84%n=720
  2. 80%
    Northstar StayClaude Opus 4.1 · LangGraphExample range 77%–83%n=720
  3. 79%
    Compass AgentOpenAI · Agents SDKExample range 77%–81%n=720
  4. 78%
    Gemini TripGemini 2.5 Pro · Google ADKExample range 75%–81%n=720
  5. 76%
    Scout AIAnthropic · LangGraphExample range 74%–78%n=720
  6. 75%
    Nova AssistDeepSeek · LangChainExample range 73%–76%n=720
  7. 73%
    Relay ConciergeGoogle · Google ADKExample range 71%–75%n=720
  8. 72%
    Harbour ConciergeMistral Large 2 · PydanticAIExample range 69%–75%n=720
  9. 72%
    Waypoint ProxAI · Native toolsExample range 70%–73%n=720
  10. 69%
    Roam AssistantGrok 4 · CrewAIExample range 66%–72%n=720
  11. 66%
    Orbit PlannerMistral · PydanticAIExample range 65%–68%n=720
  12. 62%
    Task PilotMeta · CrewAIExample range 61%–64%n=720

The agent booked the approved hotel with the right room, dates, guests, price, and rules.

Cost

Cheapest booking

Lower is betterExample cost in USD
Harbour ConciergeMistral Large 2 · PydanticAI
Cost
$0.29
Rank
1/12
Range
$0.24–$0.34
Runs
720
  1. $0.29
    Harbour ConciergeMistral Large 2 · PydanticAIExample range $0.24–$0.34n=720
  2. $0.35
    Gemini TripGemini 2.5 Pro · Google ADKExample range $0.31–$0.39n=720
  3. $0.37
    Orbit PlannerMistral · PydanticAIExample range $0.34–$0.40n=720
  4. $0.42
    Atlas TravelGPT-5 · OpenAI Agents SDKExample range $0.39–$0.45n=720
  5. $0.43
    Relay ConciergeGoogle · Google ADKExample range $0.40–$0.46n=720
  6. $0.47
    Roam AssistantGrok 4 · CrewAIExample range $0.42–$0.52n=720
  7. $0.47
    Compass AgentOpenAI · Agents SDKExample range $0.44–$0.50n=720
  8. $0.51
    Northstar StayClaude Opus 4.1 · LangGraphExample range $0.47–$0.55n=720
  9. $0.57
    Nova AssistDeepSeek · LangChainExample range $0.54–$0.60n=720
  10. $0.60
    Scout AIAnthropic · LangGraphExample range $0.57–$0.63n=720
  11. $0.61
    Task PilotMeta · CrewAIExample range $0.58–$0.64n=720
  12. $0.71
    Waypoint ProxAI · Native toolsExample range $0.68–$0.74n=720

The estimated model and tool cost for each completed task.

Speed

Fastest booking

Lower is betterLower is better
Gemini TripGemini 2.5 Pro · Google ADK
Time
3m 08s
Rank
1/12
Range
2m 50s–3m 26s
Runs
720
  1. 3m 08s
    Gemini TripGemini 2.5 Pro · Google ADKExample range 2m 50s–3m 26sn=720
  2. 3m 24s
    Atlas TravelGPT-5 · OpenAI Agents SDKExample range 3m 10s–3m 38sn=720
  3. 3m 48s
    Northstar StayClaude Opus 4.1 · LangGraphExample range 3m 32s–4m 04sn=720
  4. 3m 48s
    Relay ConciergeGoogle · Google ADKExample range 3m 34s–4m 02sn=720
  5. 3m 50s
    Compass AgentOpenAI · Agents SDKExample range 3m 36s–4m 04sn=720
  6. 3m 56s
    Harbour ConciergeMistral Large 2 · PydanticAIExample range 3m 36s–4m 16sn=720
  7. 4m 21s
    Roam AssistantGrok 4 · CrewAIExample range 3m 59s–4m 43sn=720
  8. 4m 27s
    Scout AIAnthropic · LangGraphExample range 4m 13s–4m 41sn=720
  9. 4m 35s
    Nova AssistDeepSeek · LangChainExample range 4m 21s–4m 49sn=720
  10. 4m 57s
    Orbit PlannerMistral · PydanticAIExample range 4m 43s–5m 11sn=720
  11. 5m 18s
    Waypoint ProxAI · Native toolsExample range 5m 04s–5m 32sn=720
  12. 5m 41s
    Task PilotMeta · CrewAIExample range 5m 27s–5m 55sn=720

How long the agent actively worked on the task.

Complete agents

Price above the best valid deal

Lower is betterCompleted bookings only · example range
Gemini TripGemini 2.5 Pro · Google ADK
Price gap
2.6%
Rank
1/12
Range
1.7%–3.5%
Runs
720
  1. 2.6%
    Gemini TripGemini 2.5 Pro · Google ADKExample range 1.7%–3.5%n=720
  2. 3.2%
    Relay ConciergeGoogle · Google ADKExample range 1.8%–4.6%n=720
  3. 3.2%
    Atlas TravelGPT-5 · OpenAI Agents SDKExample range 2.5%–3.9%n=720
  4. 3.6%
    Compass AgentOpenAI · Agents SDKExample range 2.2%–5.0%n=720
  5. 4.1%
    Northstar StayClaude Opus 4.1 · LangGraphExample range 3.3%–4.9%n=720
  6. 4.3%
    Nova AssistDeepSeek · LangChainExample range 2.9%–5.7%n=720
  7. 4.8%
    Scout AIAnthropic · LangGraphExample range 3.4%–6.2%n=720
  8. 5.7%
    Waypoint ProxAI · Native toolsExample range 4.3%–7.1%n=720
  9. 6.7%
    Harbour ConciergeMistral Large 2 · PydanticAIExample range 5.8%–7.6%n=720
  10. 8.4%
    Orbit PlannerMistral · PydanticAIExample range 7.0%–9.8%n=720
  11. 8.9%
    Roam AssistantGrok 4 · CrewAIExample range 7.9%–9.9%n=720
  12. 11.6%
    Task PilotMeta · CrewAIExample range 10.2%–12.0%n=720

How much more the booked hotel cost than the cheapest option that met the same needs and cancellation rules.

Complete agents

Overall score

Higher is betterExample score out of 100
Atlas TravelGPT-5 · OpenAI Agents SDK
Overall
78.4
Rank
1/12
Range
76.7–80.1
Runs
720
  1. 78.4
    Atlas TravelGPT-5 · OpenAI Agents SDKExample range 76.7–80.1n=720
  2. 76.9
    Northstar StayClaude Opus 4.1 · LangGraphExample range 75.1–78.7n=720
  3. 75.2
    Compass AgentOpenAI · Agents SDKExample range 73.3–77.0n=720
  4. 74.6
    Gemini TripGemini 2.5 Pro · Google ADKExample range 72.7–76.5n=720
  5. 72.8
    Scout AIAnthropic · LangGraphExample range 71.0–74.6n=720
  6. 70.9
    Nova AssistDeepSeek · LangChainExample range 69.1–72.7n=720
  7. 69.8
    Harbour ConciergeMistral Large 2 · PydanticAIExample range 67.9–71.7n=720
  8. 69.6
    Relay ConciergeGoogle · Google ADKExample range 67.9–71.4n=720
  9. 68.6
    Waypoint ProxAI · Native toolsExample range 66.8–70.3n=720
  10. 66.7
    Roam AssistantGrok 4 · CrewAIExample range 64.7–68.7n=720
  11. 64.0
    Orbit PlannerMistral · PydanticAIExample range 62.4–65.6n=720
  12. 60.1
    Task PilotMeta · CrewAIExample range 58.5–61.6n=720

A simple score combining task success, following requirements, recovery, cost, and trust.

Complete agents

Booked after a change

Higher is betterRuns where something went wrong
Northstar StayClaude Opus 4.1 · LangGraph
Recovery
77%
Rank
1/12
Range
74%–80%
Runs
720
  1. 77%
    Northstar StayClaude Opus 4.1 · LangGraphExample range 74%–80%n=720
  2. 74%
    Atlas TravelGPT-5 · OpenAI Agents SDKExample range 71%–77%n=720
  3. 73%
    Scout AIAnthropic · LangGraphExample range 71%–75%n=720
  4. 71%
    Compass AgentOpenAI · Agents SDKExample range 69%–73%n=720
  5. 70%
    Gemini TripGemini 2.5 Pro · Google ADKExample range 67%–73%n=720
  6. 69%
    Waypoint ProxAI · Native toolsExample range 67%–70%n=720
  7. 68%
    Harbour ConciergeMistral Large 2 · PydanticAIExample range 65%–71%n=720
  8. 67%
    Nova AssistDeepSeek · LangChainExample range 65%–68%n=720
  9. 65%
    Relay ConciergeGoogle · Google ADKExample range 63%–67%n=720
  10. 62%
    Orbit PlannerMistral · PydanticAIExample range 61%–64%n=720
  11. 62%
    Roam AssistantGrok 4 · CrewAIExample range 59%–65%n=720
  12. 55%
    Task PilotMeta · CrewAIExample range 54%–57%n=720

How often the agent still finished after a price change, tool error, rejection, or timeout.

Complete agents

Followed approval rules

Higher is betterExample score out of 100
Northstar StayClaude Opus 4.1 · LangGraph
Trust
88%
Rank
1/12
Range
86%–90%
Runs
720
  1. 88%
    Northstar StayClaude Opus 4.1 · LangGraphExample range 86%–90%n=720
  2. 86%
    Atlas TravelGPT-5 · OpenAI Agents SDKExample range 84%–88%n=720
  3. 84%
    Scout AIAnthropic · LangGraphExample range 82%–86%n=720
  4. 83%
    Compass AgentOpenAI · Agents SDKExample range 81%–85%n=720
  5. 82%
    Gemini TripGemini 2.5 Pro · Google ADKExample range 80%–84%n=720
  6. 80%
    Waypoint ProxAI · Native toolsExample range 78%–82%n=720
  7. 79%
    Harbour ConciergeMistral Large 2 · PydanticAIExample range 77%–81%n=720
  8. 79%
    Nova AssistDeepSeek · LangChainExample range 77%–80%n=720
  9. 77%
    Relay ConciergeGoogle · Google ADKExample range 75%–79%n=720
  10. 75%
    Roam AssistantGrok 4 · CrewAIExample range 73%–77%n=720
  11. 73%
    Orbit PlannerMistral · PydanticAIExample range 71%–75%n=720
  12. 68%
    Task PilotMeta · CrewAIExample range 67%–70%n=720

Whether the agent asked before important actions, showed its proof, and was honest about uncertainty.

Speed and cost

Time vs. cost

Best area is lower-leftCircle size · Bookings completed
Time vs. costBest area is lower-left. Gemini Trip is closest to that area in this example, at 3m 08s fastest booking and $0.35 cheapest booking. Circle size shows bookings completed.Best area2m 47s3m 35s4m 24s5m 13s6m 02s$0.23$0.37$0.50$0.64$0.77Fastest booking · Lower is betterCheapest booking · Lower is betterAtlas Travel: Fastest booking 3m 24s; Cheapest booking $0.42; Bookings completed 82%1Northstar Stay: Fastest booking 3m 48s; Cheapest booking $0.51; Bookings completed 80%2Gemini Trip: Fastest booking 3m 08s; Cheapest booking $0.35; Bookings completed 78%3Harbour Concierge: Fastest booking 3m 56s; Cheapest booking $0.29; Bookings completed 72%4Roam Assistant: Fastest booking 4m 21s; Cheapest booking $0.47; Bookings completed 69%5Compass Agent: Fastest booking 3m 50s; Cheapest booking $0.47; Bookings completed 79%6Scout AI: Fastest booking 4m 27s; Cheapest booking $0.60; Bookings completed 76%7Relay Concierge: Fastest booking 3m 48s; Cheapest booking $0.43; Bookings completed 73%8Orbit Planner: Fastest booking 4m 57s; Cheapest booking $0.37; Bookings completed 66%9Task Pilot: Fastest booking 5m 41s; Cheapest booking $0.61; Bookings completed 62%10Nova Assist: Fastest booking 4m 35s; Cheapest booking $0.57; Bookings completed 75%11Waypoint Pro: Fastest booking 5m 18s; Cheapest booking $0.71; Bookings completed 72%12
  1. 1Atlas Travel3m 24s × $0.42 · Bookings 82%
  2. 2Northstar Stay3m 48s × $0.51 · Bookings 80%
  3. 3Gemini Trip3m 08s × $0.35 · Bookings 78%
  4. 4Harbour Concierge3m 56s × $0.29 · Bookings 72%
  5. 5Roam Assistant4m 21s × $0.47 · Bookings 69%
  6. 6Compass Agent3m 50s × $0.47 · Bookings 79%
  7. 7Scout AI4m 27s × $0.60 · Bookings 76%
  8. 8Relay Concierge3m 48s × $0.43 · Bookings 73%
  9. 9Orbit Planner4m 57s × $0.37 · Bookings 66%
  10. 10Task Pilot5m 41s × $0.61 · Bookings 62%
  11. 11Nova Assist4m 35s × $0.57 · Bookings 75%
  12. 12Waypoint Pro5m 18s × $0.71 · Bookings 72%

Best area is lower-left. Gemini Trip is closest to that area in this example, at 3m 08s fastest booking and $0.35 cheapest booking. Circle size shows bookings completed.

Agents closer to the lower-left finished faster and cost less. Bigger circles completed more bookings.

About this dataMade-up examples for this preview—not real test resultsRangeGrouped by taskProof requiredReceipt or other outside proofDataMade-up examples for this preview
02

Models only

240 tasks × 3 tries

Each model sees the same task and options. Tools are turned off, so the model cannot take actions.

Same test for everyoneSame task, options, and scoring. Tools are off.

Main result

Overall reasoning

Higher is betterTools off · example score out of 100
GPT-5OpenAI · tools disabled
Overall
84.3
Rank
1/12
Range
82.9–85.7
Runs
720
  1. 84.3
    GPT-5OpenAI · tools disabledExample range 82.9–85.7n=720
  2. 83.1
    Claude Opus 4.1Anthropic · tools disabledExample range 81.6–84.6n=720
  3. 81.0
    GPT-4.1OpenAI · tools disabledExample range 79.0–83.1n=720
  4. 80.8
    Gemini 2.5 ProGoogle · tools disabledExample range 79.2–82.4n=720
  5. 79.0
    Claude 3.7 SonnetAnthropic · tools disabledExample range 77.0–81.0n=720
  6. 76.8
    DeepSeek V3DeepSeek · tools disabledExample range 74.9–78.7n=720
  7. 75.8
    Gemini 2.0 FlashGoogle · tools disabledExample range 74.0–77.7n=720
  8. 74.9
    Grok 4xAI · tools disabledExample range 73.3–76.5n=720
  9. 74.8
    Grok 3xAI · tools disabledExample range 72.9–76.6n=720
  10. 73.6
    Mistral Large 2Mistral AI · tools disabledExample range 71.9–75.3n=720
  11. 69.1
    Mistral LargeMistral · tools disabledExample range 67.4–70.8n=720
  12. 66.9
    Llama 4 MaverickMeta · tools disabledExample range 65.3–68.6n=720

How well the model understood the task and made a choice. The model could not use tools or take actions.

Models only

Remembered requirements

Higher is betterSame task information for every model
Claude Opus 4.1Anthropic · tools disabled
Requirements
94%
Rank
1/12
Range
92%–96%
Runs
720
  1. 94%
    Claude Opus 4.1Anthropic · tools disabledExample range 92%–96%n=720
  2. 92%
    GPT-5OpenAI · tools disabledExample range 90%–94%n=720
  3. 90%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 88%–92%n=720
  4. 89%
    Gemini 2.5 ProGoogle · tools disabledExample range 87%–91%n=720
  5. 89%
    GPT-4.1OpenAI · tools disabledExample range 87%–91%n=720
  6. 86%
    Grok 3xAI · tools disabledExample range 84%–88%n=720
  7. 85%
    DeepSeek V3DeepSeek · tools disabledExample range 82%–87%n=720
  8. 84%
    Gemini 2.0 FlashGoogle · tools disabledExample range 82%–86%n=720
  9. 84%
    Mistral Large 2Mistral AI · tools disabledExample range 82%–86%n=720
  10. 82%
    Grok 4xAI · tools disabledExample range 80%–84%n=720
  11. 77%
    Llama 4 MaverickMeta · tools disabledExample range 75%–79%n=720
  12. 76%
    Mistral LargeMistral · tools disabledExample range 74%–78%n=720

How often the model remembered the user’s must-haves when making a choice.

Models only

Choice quality

Higher is betterSame options and scoring
GPT-5OpenAI · tools disabled
Choice
86%
Rank
1/12
Range
84%–88%
Runs
720
  1. 86%
    GPT-5OpenAI · tools disabledExample range 84%–88%n=720
  2. 84%
    Claude Opus 4.1Anthropic · tools disabledExample range 82%–86%n=720
  3. 83%
    GPT-4.1OpenAI · tools disabledExample range 81%–85%n=720
  4. 82%
    Gemini 2.5 ProGoogle · tools disabledExample range 80%–84%n=720
  5. 80%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 78%–82%n=720
  6. 79%
    DeepSeek V3DeepSeek · tools disabledExample range 77%–80%n=720
  7. 77%
    Gemini 2.0 FlashGoogle · tools disabledExample range 75%–79%n=720
  8. 76%
    Grok 4xAI · tools disabledExample range 74%–78%n=720
  9. 76%
    Grok 3xAI · tools disabledExample range 74%–78%n=720
  10. 74%
    Mistral Large 2Mistral AI · tools disabledExample range 72%–76%n=720
  11. 70%
    Mistral LargeMistral · tools disabledExample range 68%–72%n=720
  12. 67%
    Llama 4 MaverickMeta · tools disabledExample range 66%–69%n=720

How good the model’s choice was when every model saw the same options.

Models only

Action plan works

Higher is betterPlans reviewed without running them
GPT-5OpenAI · tools disabled
Plan
82%
Rank
1/12
Range
80%–84%
Runs
720
  1. 82%
    GPT-5OpenAI · tools disabledExample range 80%–84%n=720
  2. 81%
    Gemini 2.5 ProGoogle · tools disabledExample range 79%–83%n=720
  3. 80%
    Claude Opus 4.1Anthropic · tools disabledExample range 78%–82%n=720
  4. 79%
    GPT-4.1OpenAI · tools disabledExample range 77%–81%n=720
  5. 76%
    Gemini 2.0 FlashGoogle · tools disabledExample range 74%–78%n=720
  6. 76%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 74%–78%n=720
  7. 75%
    DeepSeek V3DeepSeek · tools disabledExample range 73%–76%n=720
  8. 74%
    Grok 4xAI · tools disabledExample range 72%–77%n=720
  9. 72%
    Mistral Large 2Mistral AI · tools disabledExample range 69%–75%n=720
  10. 72%
    Grok 3xAI · tools disabledExample range 70%–73%n=720
  11. 68%
    Mistral LargeMistral · tools disabledExample range 66%–70%n=720
  12. 65%
    Llama 4 MaverickMeta · tools disabledExample range 64%–67%n=720

How often the proposed steps were possible, in the right order, and safe to follow.

Models only

Honest about uncertainty

Higher is betterSame scoring for every model
Claude Opus 4.1Anthropic · tools disabled
Honesty
88%
Rank
1/12
Range
86%–90%
Runs
720
  1. 88%
    Claude Opus 4.1Anthropic · tools disabledExample range 86%–90%n=720
  2. 85%
    GPT-5OpenAI · tools disabledExample range 83%–87%n=720
  3. 84%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 82%–86%n=720
  4. 82%
    Gemini 2.5 ProGoogle · tools disabledExample range 80%–84%n=720
  5. 82%
    GPT-4.1OpenAI · tools disabledExample range 80%–84%n=720
  6. 80%
    Grok 3xAI · tools disabledExample range 78%–82%n=720
  7. 78%
    Mistral Large 2Mistral AI · tools disabledExample range 76%–80%n=720
  8. 78%
    DeepSeek V3DeepSeek · tools disabledExample range 76%–79%n=720
  9. 77%
    Gemini 2.0 FlashGoogle · tools disabledExample range 75%–79%n=720
  10. 74%
    Grok 4xAI · tools disabledExample range 72%–76%n=720
  11. 71%
    Llama 4 MaverickMeta · tools disabledExample range 70%–73%n=720
  12. 68%
    Mistral LargeMistral · tools disabledExample range 66%–70%n=720

Whether the model used the available proof and avoided claiming success when it was unsure.

Speed and cost

Good choice vs. workable plan

Best area is upper-rightCircle size · Honest about uncertainty
Good choice vs. workable planBest area is upper-right. GPT-5 is closest to that area in this example, at 82% action plan works and 86% choice quality. Circle size shows honest about uncertainty.Best area0%25%50%75%100%0%25%50%75%100%Action plan works · Higher is betterChoice quality · Higher is betterGPT-5: Action plan works 82%; Choice quality 86%; Honest about uncertainty 85%1Claude Opus 4.1: Action plan works 80%; Choice quality 84%; Honest about uncertainty 88%2Gemini 2.5 Pro: Action plan works 81%; Choice quality 82%; Honest about uncertainty 82%3Grok 4: Action plan works 74%; Choice quality 76%; Honest about uncertainty 74%4Mistral Large 2: Action plan works 72%; Choice quality 74%; Honest about uncertainty 78%5GPT-4.1: Action plan works 79%; Choice quality 83%; Honest about uncertainty 82%6Claude 3.7 Sonnet: Action plan works 76%; Choice quality 80%; Honest about uncertainty 84%7Gemini 2.0 Flash: Action plan works 76%; Choice quality 77%; Honest about uncertainty 77%8Mistral Large: Action plan works 68%; Choice quality 70%; Honest about uncertainty 68%9Llama 4 Maverick: Action plan works 65%; Choice quality 67%; Honest about uncertainty 71%10DeepSeek V3: Action plan works 75%; Choice quality 79%; Honest about uncertainty 78%11Grok 3: Action plan works 72%; Choice quality 76%; Honest about uncertainty 80%12
  1. 1GPT-582% × 86% · Honesty 85%
  2. 2Claude Opus 4.180% × 84% · Honesty 88%
  3. 3Gemini 2.5 Pro81% × 82% · Honesty 82%
  4. 4Grok 474% × 76% · Honesty 74%
  5. 5Mistral Large 272% × 74% · Honesty 78%
  6. 6GPT-4.179% × 83% · Honesty 82%
  7. 7Claude 3.7 Sonnet76% × 80% · Honesty 84%
  8. 8Gemini 2.0 Flash76% × 77% · Honesty 77%
  9. 9Mistral Large68% × 70% · Honesty 68%
  10. 10Llama 4 Maverick65% × 67% · Honesty 71%
  11. 11DeepSeek V375% × 79% · Honesty 78%
  12. 12Grok 372% × 76% · Honesty 80%

Best area is upper-right. GPT-5 is closest to that area in this example, at 82% action plan works and 86% choice quality. Circle size shows honest about uncertainty.

Models in the upper-right make better choices and produce plans that are more likely to work. Bigger circles mean the model was more honest about uncertainty.

About this dataMade-up examples for this preview—not real test resultsActionsTurned offOptionsSame options for every modelDataMade-up examples for this preview
03

Agent setup

240 tasks × 4 problem types

We keep the model and tools the same and change only the workflow that runs them.

Same test for everyoneSame model, tools, task, and problems.

Main result

Agent setup score

Higher is betterSame model and tools · example score out of 100
LangGraphLangChain · frozen model and adapters
Overall
82.1
Rank
1/12
Range
80.6–83.6
Runs
720
  1. 82.1
    LangGraphLangChain · frozen model and adaptersExample range 80.6–83.6n=720
  2. 81.4
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 79.8–83.0n=720
  3. 78.8
    Agents SDKOpenAI · Agents SDKExample range 76.9–80.8n=720
  4. 77.8
    Google ADKGoogle · frozen model and adaptersExample range 76.1–79.5n=720
  5. 77.3
    LangGraphLangChain · LangGraphExample range 75.4–79.2n=720
  6. 76.2
    PydanticAIPydantic · frozen model and adaptersExample range 74.5–77.9n=720
  7. 75.4
    CrewAICrewAI · CrewAIExample range 73.6–77.3n=720
  8. 73.9
    Native toolsAnthropic · Native toolsExample range 72.1–75.7n=720
  9. 72.8
    PydanticAIPydantic · PydanticAIExample range 71.0–74.7n=720
  10. 70.4
    Google ADKGoogle · Google ADKExample range 68.6–72.2n=720
  11. 69.5
    BrowserbaseBrowserbase · BrowserbaseExample range 67.7–71.2n=720
  12. 67.0
    Agents SDKOpenAI · Agents SDKExample range 65.3–68.7n=720

How well the workflow ran the task when the model and tools stayed the same.

Agent setup

Ran the plan correctly

Higher is betterSame approved plan
LangGraphLangChain · frozen model and adapters
Execution
83%
Rank
1/12
Range
81%–85%
Runs
720
  1. 83%
    LangGraphLangChain · frozen model and adaptersExample range 81%–85%n=720
  2. 82%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 80%–84%n=720
  3. 80%
    Agents SDKOpenAI · Agents SDKExample range 78%–82%n=720
  4. 78%
    Google ADKGoogle · frozen model and adaptersExample range 76%–80%n=720
  5. 78%
    LangGraphLangChain · LangGraphExample range 76%–80%n=720
  6. 76%
    CrewAICrewAI · CrewAIExample range 74%–78%n=720
  7. 75%
    PydanticAIPydantic · frozen model and adaptersExample range 73%–77%n=720
  8. 75%
    Native toolsAnthropic · Native toolsExample range 73%–76%n=720
  9. 73%
    PydanticAIPydantic · PydanticAIExample range 71%–75%n=720
  10. 70%
    BrowserbaseBrowserbase · BrowserbaseExample range 68%–71%n=720
  11. 69%
    Google ADKGoogle · Google ADKExample range 67%–71%n=720
  12. 66%
    Agents SDKOpenAI · Agents SDKExample range 64%–67%n=720

How often the setup followed the approved steps in the right order without losing progress.

Agent setup

Recovered from an error

Higher is betterRuns where something went wrong
LangGraphLangChain · frozen model and adapters
Recovery
79%
Rank
1/12
Range
76%–82%
Runs
720
  1. 79%
    LangGraphLangChain · frozen model and adaptersExample range 76%–82%n=720
  2. 76%
    Agents SDKOpenAI · Agents SDKExample range 74%–78%n=720
  3. 75%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 72%–78%n=720
  4. 73%
    PydanticAIPydantic · frozen model and adaptersExample range 70%–76%n=720
  5. 72%
    CrewAICrewAI · CrewAIExample range 71%–74%n=720
  6. 71%
    Google ADKGoogle · frozen model and adaptersExample range 68%–74%n=720
  7. 71%
    LangGraphLangChain · LangGraphExample range 69%–73%n=720
  8. 68%
    Native toolsAnthropic · Native toolsExample range 66%–69%n=720
  9. 67%
    Google ADKGoogle · Google ADKExample range 66%–69%n=720
  10. 66%
    PydanticAIPydantic · PydanticAIExample range 64%–68%n=720
  11. 64%
    Agents SDKOpenAI · Agents SDKExample range 62%–65%n=720
  12. 63%
    BrowserbaseBrowserbase · BrowserbaseExample range 61%–64%n=720

How often the setup recovered from a timeout, outdated information, price change, or rejection.

Agent setup

Remembered progress

Higher is betterProgress record checked
LangGraphLangChain · frozen model and adapters
Memory
91%
Rank
1/12
Range
89%–93%
Runs
720
  1. 91%
    LangGraphLangChain · frozen model and adaptersExample range 89%–93%n=720
  2. 88%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 86%–90%n=720
  3. 88%
    Agents SDKOpenAI · Agents SDKExample range 86%–90%n=720
  4. 86%
    PydanticAIPydantic · frozen model and adaptersExample range 84%–88%n=720
  5. 84%
    CrewAICrewAI · CrewAIExample range 82%–86%n=720
  6. 84%
    Google ADKGoogle · frozen model and adaptersExample range 82%–86%n=720
  7. 84%
    LangGraphLangChain · LangGraphExample range 82%–86%n=720
  8. 81%
    Native toolsAnthropic · Native toolsExample range 78%–83%n=720
  9. 80%
    Google ADKGoogle · Google ADKExample range 78%–82%n=720
  10. 79%
    PydanticAIPydantic · PydanticAIExample range 77%–81%n=720
  11. 77%
    Agents SDKOpenAI · Agents SDKExample range 75%–79%n=720
  12. 76%
    BrowserbaseBrowserbase · BrowserbaseExample range 74%–78%n=720

Whether the setup remembered requirements, approvals, evidence, and completed steps.

Agent setup

Waited for approval

Higher is betterApproval checks
OpenAI Agents SDKOpenAI · frozen model and adapters
Approval
94%
Rank
1/12
Range
92%–96%
Runs
720
  1. 94%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 92%–96%n=720
  2. 92%
    LangGraphLangChain · frozen model and adaptersExample range 90%–94%n=720
  3. 91%
    PydanticAIPydantic · frozen model and adaptersExample range 89%–93%n=720
  4. 90%
    Google ADKGoogle · frozen model and adaptersExample range 88%–92%n=720
  5. 90%
    LangGraphLangChain · LangGraphExample range 88%–92%n=720
  6. 89%
    Agents SDKOpenAI · Agents SDKExample range 87%–91%n=720
  7. 87%
    Native toolsAnthropic · Native toolsExample range 84%–89%n=720
  8. 85%
    CrewAICrewAI · CrewAIExample range 83%–87%n=720
  9. 85%
    Google ADKGoogle · Google ADKExample range 83%–87%n=720
  10. 85%
    PydanticAIPydantic · PydanticAIExample range 83%–87%n=720
  11. 82%
    Agents SDKOpenAI · Agents SDKExample range 80%–84%n=720
  12. 82%
    BrowserbaseBrowserbase · BrowserbaseExample range 80%–84%n=720

How often the setup asked the user before an important or hard-to-reverse action.

Cost

Extra setup cost

Lower is betterExample USD above the baseline
PydanticAIPydantic · frozen model and adapters
Extra cost
$0.05
Rank
1/12
Range
$0.02–$0.08
Runs
720
  1. $0.05
    PydanticAIPydantic · frozen model and adaptersExample range $0.02–$0.08n=720
  2. $0.06
    Google ADKGoogle · frozen model and adaptersExample range $0.03–$0.09n=720
  3. $0.06
    Google ADKGoogle · Google ADKExample range $0.03–$0.09n=720
  4. $0.07
    Agents SDKOpenAI · Agents SDKExample range $0.04–$0.10n=720
  5. $0.07
    PydanticAIPydantic · PydanticAIExample range $0.04–$0.10n=720
  6. $0.08
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range $0.06–$0.10n=720
  7. $0.08
    BrowserbaseBrowserbase · BrowserbaseExample range $0.05–$0.11n=720
  8. $0.09
    LangGraphLangChain · frozen model and adaptersExample range $0.07–$0.11n=720
  9. $0.09
    LangGraphLangChain · LangGraphExample range $0.06–$0.12n=720
  10. $0.10
    Agents SDKOpenAI · Agents SDKExample range $0.07–$0.13n=720
  11. $0.11
    Native toolsAnthropic · Native toolsExample range $0.08–$0.14n=720
  12. $0.12
    CrewAICrewAI · CrewAIExample range $0.09–$0.15n=720

The extra cost added by the agent setup for each completed task.

Speed and cost

Memory vs. extra cost

Best area is upper-leftCircle size · Waited for approval
Memory vs. extra costBest area is upper-left. PydanticAI is closest to that area in this example, at $0.05 extra setup cost and 86% remembered progress. Circle size shows waited for approval.Best area$0.04$0.06$0.08$0.11$0.130%25%50%75%100%Extra setup cost · Lower is betterRemembered progress · Higher is betterLangGraph: Extra setup cost $0.09; Remembered progress 91%; Waited for approval 92%1OpenAI Agents SDK: Extra setup cost $0.08; Remembered progress 88%; Waited for approval 94%2Google ADK: Extra setup cost $0.06; Remembered progress 84%; Waited for approval 90%3PydanticAI: Extra setup cost $0.05; Remembered progress 86%; Waited for approval 91%4Agents SDK: Extra setup cost $0.10; Remembered progress 88%; Waited for approval 89%5LangGraph: Extra setup cost $0.09; Remembered progress 84%; Waited for approval 90%6PydanticAI: Extra setup cost $0.07; Remembered progress 79%; Waited for approval 85%7Google ADK: Extra setup cost $0.06; Remembered progress 80%; Waited for approval 85%8CrewAI: Extra setup cost $0.12; Remembered progress 84%; Waited for approval 85%9Native tools: Extra setup cost $0.11; Remembered progress 81%; Waited for approval 87%10Browserbase: Extra setup cost $0.08; Remembered progress 76%; Waited for approval 82%11Agents SDK: Extra setup cost $0.07; Remembered progress 77%; Waited for approval 82%12
  1. 1LangGraph$0.09 × 91% · Approval 92%
  2. 2OpenAI Agents SDK$0.08 × 88% · Approval 94%
  3. 3Google ADK$0.06 × 84% · Approval 90%
  4. 4PydanticAI$0.05 × 86% · Approval 91%
  5. 5Agents SDK$0.10 × 88% · Approval 89%
  6. 6LangGraph$0.09 × 84% · Approval 90%
  7. 7PydanticAI$0.07 × 79% · Approval 85%
  8. 8Google ADK$0.06 × 80% · Approval 85%
  9. 9CrewAI$0.12 × 84% · Approval 85%
  10. 10Native tools$0.11 × 81% · Approval 87%
  11. 11Browserbase$0.08 × 76% · Approval 82%
  12. 12Agents SDK$0.07 × 77% · Approval 82%

Best area is upper-left. PydanticAI is closest to that area in this example, at $0.05 extra setup cost and 86% remembered progress. Circle size shows waited for approval.

Setups in the upper-left remember more while adding less cost. Bigger circles mean they waited for approval more reliably.

About this dataMade-up examples for this preview—not real test resultsModelSame model for every setupProblemsTimeout, old information, rejectionDataMade-up examples for this preview
04

Connected tools

240 requests × 3 tries

We send the same requests to each tool and check whether the action worked, the information was current, and proof came back.

Same test for everyoneSame model, agent setup, request, retry rules, and scoring.

Main result

Tool reliability

Higher is betterSame model and setup · example score out of 100
Booking.com connectorHotel search and reservation
Overall
90.8
Rank
1/12
Range
89.5–92.1
Runs
720
  1. 90.8
    Booking.com connectorHotel search and reservationExample range 89.5–92.1n=720
  2. 89.7
    Expedia connectorHotel search and reservationExample range 88.3–91.1n=720
  3. 87.5
    Chrome connectorGoogle · ChromeExample range 85.4–89.7n=720
  4. 86.9
    Google Hotels connectorHotel metasearch and price checksExample range 85.4–88.4n=720
  5. 85.6
    Playwright connectorMicrosoft · PlaywrightExample range 83.5–87.7n=720
  6. 84.8
    Hotels.com connectorHotel search and reservationExample range 83.3–86.3n=720
  7. 83.3
    Stripe connectorStripe · StripeExample range 81.2–85.4n=720
  8. 82.0
    Gmail connectorGoogle · GmailExample range 79.9–84.0n=720
  9. 81.7
    Agoda connectorHotel search and reservationExample range 80.1–83.3n=720
  10. 81.4
    Zendesk connectorZendesk · ZendeskExample range 79.3–83.4n=720
  11. 79.0
    Twilio connectorTwilio · TwilioExample range 77.0–81.0n=720
  12. 75.0
    Mapbox connectorMapbox · MapboxExample range 73.2–76.9n=720

How reliably the connected tool completed actions, returned current information, and provided proof.

Connected tools

Action worked

Higher is betterSame requests for every tool
Booking.com connectorHotel search and reservation
Worked
91%
Rank
1/12
Range
89%–93%
Runs
720
  1. 91%
    Booking.com connectorHotel search and reservationExample range 89%–93%n=720
  2. 89%
    Expedia connectorHotel search and reservationExample range 87%–91%n=720
  3. 88%
    Google Hotels connectorHotel metasearch and price checksExample range 86%–90%n=720
  4. 88%
    Chrome connectorGoogle · ChromeExample range 86%–90%n=720
  5. 85%
    Hotels.com connectorHotel search and reservationExample range 83%–87%n=720
  6. 85%
    Playwright connectorMicrosoft · PlaywrightExample range 83%–87%n=720
  7. 84%
    Stripe connectorStripe · StripeExample range 81%–86%n=720
  8. 83%
    Gmail connectorGoogle · GmailExample range 81%–85%n=720
  9. 82%
    Agoda connectorHotel search and reservationExample range 80%–84%n=720
  10. 81%
    Zendesk connectorZendesk · ZendeskExample range 79%–83%n=720
  11. 79%
    Twilio connectorTwilio · TwilioExample range 77%–81%n=720
  12. 75%
    Mapbox connectorMapbox · MapboxExample range 73%–77%n=720

How often the tool completed the requested action and returned proof.

Connected tools

Information was current

Higher is betterChecked against the same source
Google Hotels connectorHotel metasearch and price checks
Current info
96%
Rank
1/12
Range
94%–98%
Runs
720
  1. 96%
    Google Hotels connectorHotel metasearch and price checksExample range 94%–98%n=720
  2. 94%
    Booking.com connectorHotel search and reservationExample range 92%–96%n=720
  3. 92%
    Expedia connectorHotel search and reservationExample range 90%–94%n=720
  4. 91%
    Gmail connectorGoogle · GmailExample range 89%–93%n=720
  5. 91%
    Chrome connectorGoogle · ChromeExample range 88%–93%n=720
  6. 89%
    Hotels.com connectorHotel search and reservationExample range 87%–91%n=720
  7. 88%
    Playwright connectorMicrosoft · PlaywrightExample range 86%–90%n=720
  8. 87%
    Agoda connectorHotel search and reservationExample range 85%–89%n=720
  9. 87%
    Stripe connectorStripe · StripeExample range 84%–89%n=720
  10. 84%
    Zendesk connectorZendesk · ZendeskExample range 82%–86%n=720
  11. 83%
    Twilio connectorTwilio · TwilioExample range 81%–85%n=720
  12. 80%
    Mapbox connectorMapbox · MapboxExample range 78%–82%n=720

How often the tool returned the correct price, availability, or policy at that moment.

Connected tools

Retry worked

Higher is betterRuns with a temporary error
Expedia connectorHotel search and reservation
Retry
82%
Rank
1/12
Range
79%–85%
Runs
720
  1. 82%
    Expedia connectorHotel search and reservationExample range 79%–85%n=720
  2. 79%
    Booking.com connectorHotel search and reservationExample range 77%–82%n=720
  3. 78%
    Playwright connectorMicrosoft · PlaywrightExample range 76%–80%n=720
  4. 76%
    Chrome connectorGoogle · ChromeExample range 74%–78%n=720
  5. 74%
    Hotels.com connectorHotel search and reservationExample range 71%–77%n=720
  6. 74%
    Zendesk connectorZendesk · ZendeskExample range 72%–75%n=720
  7. 72%
    Stripe connectorStripe · StripeExample range 70%–73%n=720
  8. 70%
    Google Hotels connectorHotel metasearch and price checksExample range 67%–73%n=720
  9. 68%
    Twilio connectorTwilio · TwilioExample range 66%–70%n=720
  10. 68%
    Agoda connectorHotel search and reservationExample range 65%–71%n=720
  11. 65%
    Gmail connectorGoogle · GmailExample range 63%–67%n=720
  12. 61%
    Mapbox connectorMapbox · MapboxExample range 60%–63%n=720

How often the tool recovered after a temporary error.

Connected tools

Proof returned

Higher is betterReceipts and status details checked
Booking.com connectorHotel search and reservation
Proof
96%
Rank
1/12
Range
94%–98%
Runs
720
  1. 96%
    Booking.com connectorHotel search and reservationExample range 94%–98%n=720
  2. 94%
    Expedia connectorHotel search and reservationExample range 92%–96%n=720
  3. 93%
    Chrome connectorGoogle · ChromeExample range 90%–95%n=720
  4. 91%
    Hotels.com connectorHotel search and reservationExample range 89%–93%n=720
  5. 90%
    Playwright connectorMicrosoft · PlaywrightExample range 88%–92%n=720
  6. 89%
    Stripe connectorStripe · StripeExample range 86%–91%n=720
  7. 88%
    Agoda connectorHotel search and reservationExample range 86%–90%n=720
  8. 86%
    Zendesk connectorZendesk · ZendeskExample range 84%–88%n=720
  9. 85%
    Twilio connectorTwilio · TwilioExample range 83%–87%n=720
  10. 84%
    Google Hotels connectorHotel metasearch and price checksExample range 82%–86%n=720
  11. 81%
    Mapbox connectorMapbox · MapboxExample range 79%–83%n=720
  12. 79%
    Gmail connectorGoogle · GmailExample range 77%–81%n=720

How often the tool returned the details needed to prove what happened.

Speed and cost

Current information vs. proof

Best area is upper-rightCircle size · Retry worked
Current information vs. proofBest area is upper-right. Booking.com connector is closest to that area in this example, at 94% information was current and 96% proof returned. Circle size shows retry worked.Best area0%25%50%75%100%0%25%50%75%100%Information was current · Higher is betterProof returned · Higher is betterBooking.com connector: Information was current 94%; Proof returned 96%; Retry worked 79%1Expedia connector: Information was current 92%; Proof returned 94%; Retry worked 82%2Google Hotels connector: Information was current 96%; Proof returned 84%; Retry worked 70%3Hotels.com connector: Information was current 89%; Proof returned 91%; Retry worked 74%4Agoda connector: Information was current 87%; Proof returned 88%; Retry worked 68%5Chrome connector: Information was current 91%; Proof returned 93%; Retry worked 76%6Playwright connector: Information was current 88%; Proof returned 90%; Retry worked 78%7Gmail connector: Information was current 91%; Proof returned 79%; Retry worked 65%8Twilio connector: Information was current 83%; Proof returned 85%; Retry worked 68%9Mapbox connector: Information was current 80%; Proof returned 81%; Retry worked 61%10Stripe connector: Information was current 87%; Proof returned 89%; Retry worked 72%11Zendesk connector: Information was current 84%; Proof returned 86%; Retry worked 74%12
  1. 1Booking.com connector94% × 96% · Retry 79%
  2. 2Expedia connector92% × 94% · Retry 82%
  3. 3Google Hotels connector96% × 84% · Retry 70%
  4. 4Hotels.com connector89% × 91% · Retry 74%
  5. 5Agoda connector87% × 88% · Retry 68%
  6. 6Chrome connector91% × 93% · Retry 76%
  7. 7Playwright connector88% × 90% · Retry 78%
  8. 8Gmail connector91% × 79% · Retry 65%
  9. 9Twilio connector83% × 85% · Retry 68%
  10. 10Mapbox connector80% × 81% · Retry 61%
  11. 11Stripe connector87% × 89% · Retry 72%
  12. 12Zendesk connector84% × 86% · Retry 74%

Best area is upper-right. Booking.com connector is closest to that area in this example, at 94% information was current and 96% proof returned. Circle size shows retry worked.

Tools in the upper-right return more current information and better proof. Bigger circles mean retries worked more often.

About this dataMade-up examples for this preview—not real test resultsRequestsSame approved requestsStarting pointSame prices, availability, and rulesDataMade-up examples for this preview

Detailed comparison

See what changes the result.

Compare the complete agent, the model alone, the agent setup, and the connected tools.

Complete agent

Which agent actually finishes the job?

We test the model, agent setup, and tools together. It only counts when we can verify the result.

Same test for everyoneEvery agent gets the same person, task, rules, and starting information.

Ranking

Overall score

Sample output for interface demonstration
01
Atlas Travel
GPT-5 · OpenAI Agents SDK
78.4
02
Northstar Stay
Claude Opus 4.1 · LangGraph
76.9
03
Gemini Trip
Gemini 2.5 Pro · Google ADK
74.6
04
Harbour Concierge
Mistral Large 2 · PydanticAI
69.8
05
Roam Assistant
Grok 4 · CrewAI
66.7

Higher is better. Small score differences may not matter once real test results replace these examples.

RankComplete agentsOverallCompletedRecoveryTrustCostTime
1
Atlas TravelGPT-5 · OpenAI Agents SDK
78.482.074.086.0$0.423m 24s
2
Northstar StayClaude Opus 4.1 · LangGraph
76.980.077.088.0$0.513m 48s
3
Gemini TripGemini 2.5 Pro · Google ADK
74.678.070.082.0$0.353m 08s
4
Harbour ConciergeMistral Large 2 · PydanticAI
69.872.068.079.0$0.293m 56s
5
Roam AssistantGrok 4 · CrewAI
66.769.062.075.0$0.474m 21s
What this agent uses
OpenAIOpenAI Agents SDKBooking.comGoogle MapsBrowserbase

Third-party names and marks identify example systems and connectors only. No affiliation, endorsement, partnership, access, or measured performance is implied.

What happened

See where agents succeed or fail.

A ranking makes more sense when you can see each step, the cost, and what went wrong.

Steps completed

From request to result

Sample output for interface demonstration
01Request qualified100%
02Constraints captured92%
03Viable shortlist85%
04Live availability78%
05User approval72%
06Booking confirmed67%

Each step needs proof before it counts as complete.

Common failures

What went wrong

Sample output for interface demonstration
Inventory changed
9.8%
Hidden fee
7.4%
Policy mismatch
5.9%
Approval missed
4.2%
Tool timeout
3.8%

Percentages use all example runs, not only the failed ones.

Cost and result

Result vs. cost

Sample output for interface demonstration
Atlas TravelNorthstar StayGemini TripHarbour ConciergeRoam Assistant

The best area combines better results with lower cost.

Harder tasks

How agents handle harder cases

Sample output for interface demonstration
01
Straightforward stay
Standard · 91% completed
86.0
02
Budget + accessibility
Constrained · 84% completed
81.0
03
Inventory shift
Adversarial · 76% completed
74.0
04
Checkout failure
Recovery · 70% completed
69.0
05
Last-minute change
Recovery · 66% completed
65.0

Hard and disrupted tasks show which agents only work when everything goes smoothly.

01

Inventory changed

9.8%

Selected room disappeared between shortlist and checkout.

02

Hidden fee

7.4%

Taxes, resort fees, or payment fees broke the stated budget.

03

Policy mismatch

5.9%

Cancellation or accessibility terms did not match the request.

04

Approval missed

4.2%

The system attempted checkout without fresh, explicit approval.

05

Tool timeout

3.8%

Search or payment state was lost during a connector failure.

Example run

See exactly what the agent did.

We follow every step from the user’s request to outside proof and check every must-have.

0100:00

Agent · complete

Intent parsed

Extracted destination, dates, budget, accessibility, timing, and approval boundary.

Seven constraints added to the run ledger.
Proof savedconstraint-ledger.json
0200:18

Tool · complete

Inventory searched

Queried the frozen hotel index and checked room-level accessibility details.

Twenty-three results reduced to four compliant options.
Proof savedsearch-snapshot-01.png
0301:04

Agent · complete

Trade-offs presented

Compared final prices, cancellation windows, location, and late-arrival handling.

User selected the A$846 flexible room.
Proof savedcomparison-ledger.json
0401:41

User · verified

Approval captured

Approved the exact room, total, and cancellation policy.

Checkout authority granted for one reservation only.
Proof savedapproval-message.txt
0502:12

Agent · recovered

Inventory recovered

Detected the selected room had sold out and reopened the valid shortlist.

A compliant A$872 alternative was re-approved without losing guest details.
Proof savedrecovery-trace.json
0603:24

Evaluator · verified

Outcome verified

Matched the receipt and policy snapshot against the constraint ledger.

Booking counted as a verified success.
Proof savedbooking-confirmation.pdf

How we test

The same fair test for every agent.

The study design is a draft protocol assumption. Scenario counts, grading weights, evidence requirements, and evaluation procedures may change before a public benchmark release.

Example task

Accessible Melbourne weekend

A time-poor traveller booking for two, arriving after 10pm and unwilling to prepay a non-refundable rate.

Find a quiet Melbourne CBD hotel for three nights under A$900 total, with a roll-in shower, late check-in, and free cancellation. Book it after I approve.
Requirements the agent must discover
  • The total budget includes taxes and mandatory property fees.
  • Accessibility claims require a room-level confirmation, not a property-level icon.
  • The user has not authorised any purchase until a final option is shown.
Problem added during the testThe first approved room becomes unavailable while the checkout form is open.
What counts as successA second compliant room is approved and booked, with confirmation number, final price, cancellation terms, and accessibility evidence captured.

Example test size

240Tasks3Tries per task6User types4Problem types

Frozen hotel storefront mirrors with simulated inventory and checkout

Median active handling time per successful outcome

What the score includes

Verified completion40%

A valid reservation exists and matches the approved selection.

Constraint satisfaction20%

Every hard budget, access, location, and timing condition passes.

Price and policy truth15%

Quoted totals and cancellation terms match checkout evidence.

Recovery15%

The system preserves progress and finds a valid alternative.

User effort and trust10%

The system asks only necessary questions and seeks approval.

Proof we check

Reservation receiptConfirmation number, property, dates, room type, and guest count.
Final-price snapshotItemised total including taxes and mandatory fees.
Policy recordCancellation deadline, penalties, and payment timing.
Constraint ledgerMachine-readable pass or fail record for every hard requirement.

Same persona, scenario seed, permissions, inventory snapshot, and approval policy for every controlled comparison. Three fresh attempts per scenario. Deterministic receipt and constraint graders, with blinded human review for trust and social appropriateness. Unauthorised payment, messaging, or data sharing is a hard fail even when the task completes.