Agent tests/FLT-03/Preview

Real-world task test · FLT-03

Flight Booking Reliability

Can an agent book the right flight after bags, fees, and connection risks are counted?

We test the full price, bags, airports, connections, refund rules, user preferences, booking, and recovery when a flight changes.

Highlights · FLT-03

Made-up example data
01

Complete agents

300 tasks × 3 tries = 900 example runs

We test the model, agent setup, and tools together. It only counts when we can verify the result.

Same test for everyoneEvery agent gets the same person, task, rules, and starting information.

Main result

Correct flights booked

Higher is betterBased on the example runs
RoutekeeperGPT-5 · OpenAI Agents SDK
Booked
71%
Rank
1/12
Range
69%–73%
Runs
900
  1. 71%
    RoutekeeperGPT-5 · OpenAI Agents SDKExample range 69%–73%n=900
  2. 69%
    Aero GuideClaude Opus 4.1 · LangGraphExample range 66%–72%n=900
  3. 68%
    Flight LensGemini 2.5 Pro · Google ADKExample range 65%–71%n=900
  4. 68%
    Compass AgentOpenAI · Agents SDKExample range 66%–69%n=900
  5. 65%
    Scout AIAnthropic · LangGraphExample range 63%–67%n=900
  6. 64%
    Nova AssistDeepSeek · LangChainExample range 62%–65%n=900
  7. 63%
    Relay ConciergeGoogle · Google ADKExample range 61%–65%n=900
  8. 63%
    Open SkiesMistral Large 2 · PydanticAIExample range 60%–66%n=900
  9. 61%
    Waypoint ProxAI · Native toolsExample range 59%–62%n=900
  10. 59%
    Waypoint AirGrok 4 · CrewAIExample range 56%–62%n=900
  11. 57%
    Orbit PlannerMistral · PydanticAIExample range 56%–59%n=900
  12. 52%
    Task PilotMeta · CrewAIExample range 51%–54%n=900

The ticket matched the route, dates, bags, airports, connection rules, price, and user approval.

Cost

Cheapest booking

Lower is betterExample cost in USD
Open SkiesMistral Large 2 · PydanticAI
Cost
$0.74
Rank
1/12
Range
$0.69–$0.79
Runs
900
  1. $0.74
    Open SkiesMistral Large 2 · PydanticAIExample range $0.69–$0.79n=900
  2. $0.88
    Flight LensGemini 2.5 Pro · Google ADKExample range $0.84–$0.92n=900
  3. $0.93
    Orbit PlannerMistral · PydanticAIExample range $0.90–$0.96n=900
  4. $1.07
    Relay ConciergeGoogle · Google ADKExample range $1.04–$1.10n=900
  5. $1.08
    Waypoint AirGrok 4 · CrewAIExample range $1.03–$1.13n=900
  6. $1.14
    RoutekeeperGPT-5 · OpenAI Agents SDKExample range $1.11–$1.17n=900
  7. $1.28
    Compass AgentOpenAI · Agents SDKExample range $1.25–$1.31n=900
  8. $1.32
    Aero GuideClaude Opus 4.1 · LangGraphExample range $1.28–$1.36n=900
  9. $1.41
    Task PilotMeta · CrewAIExample range $1.38–$1.44n=900
  10. $1.54
    Nova AssistDeepSeek · LangChainExample range $1.51–$1.57n=900
  11. $1.54
    Scout AIAnthropic · LangGraphExample range $1.51–$1.57n=900
  12. $1.84
    Waypoint ProxAI · Native toolsExample range $1.81–$1.87n=900

The estimated model and tool cost for each completed task.

Speed

Fastest booking

Lower is betterLower is better
Flight LensGemini 2.5 Pro · Google ADK
Time
7m 16s
Rank
1/12
Range
6m 58s–7m 34s
Runs
900
  1. 7m 16s
    Flight LensGemini 2.5 Pro · Google ADKExample range 6m 58s–7m 34sn=900
  2. 7m 48s
    RoutekeeperGPT-5 · OpenAI Agents SDKExample range 7m 34s–8m 02sn=900
  3. 8m 32s
    Aero GuideClaude Opus 4.1 · LangGraphExample range 8m 16s–8m 48sn=900
  4. 8m 47s
    Compass AgentOpenAI · Agents SDKExample range 8m 33s–9m 01sn=900
  5. 8m 50s
    Relay ConciergeGoogle · Google ADKExample range 8m 36s–9m 04sn=900
  6. 9m 08s
    Open SkiesMistral Large 2 · PydanticAIExample range 8m 48s–9m 28sn=900
  7. 9m 59s
    Scout AIAnthropic · LangGraphExample range 9m 45s–10m 13sn=900
  8. 10m 01s
    Waypoint AirGrok 4 · CrewAIExample range 9m 39s–10m 23sn=900
  9. 10m 32s
    Nova AssistDeepSeek · LangChainExample range 10m 18s–10m 46sn=900
  10. 11m 30s
    Orbit PlannerMistral · PydanticAIExample range 11m 16s–11m 44sn=900
  11. 11m 54s
    Waypoint ProxAI · Native toolsExample range 11m 40s–12m 08sn=900
  12. 13m 04s
    Task PilotMeta · CrewAIExample range 12m 50s–13m 18sn=900

How long the agent actively worked on the task.

Complete agents

Price above the best valid flight

Lower is betterCompleted bookings only · example range
Flight LensGemini 2.5 Pro · Google ADK
Price gap
3.5%
Rank
1/12
Range
2.6%–4.4%
Runs
900
  1. 3.5%
    Flight LensGemini 2.5 Pro · Google ADKExample range 2.6%–4.4%n=900
  2. 4.3%
    Relay ConciergeGoogle · Google ADKExample range 2.9%–5.7%n=900
  3. 4.8%
    RoutekeeperGPT-5 · OpenAI Agents SDKExample range 4.1%–5.5%n=900
  4. 5.4%
    Compass AgentOpenAI · Agents SDKExample range 4.0%–6.8%n=900
  5. 5.6%
    Aero GuideClaude Opus 4.1 · LangGraphExample range 4.8%–6.4%n=900
  6. 6.5%
    Nova AssistDeepSeek · LangChainExample range 5.1%–7.9%n=900
  7. 6.6%
    Scout AIAnthropic · LangGraphExample range 5.2%–8.0%n=900
  8. 7.4%
    Open SkiesMistral Large 2 · PydanticAIExample range 6.5%–8.3%n=900
  9. 7.8%
    Waypoint ProxAI · Native toolsExample range 6.4%–9.2%n=900
  10. 9.3%
    Orbit PlannerMistral · PydanticAIExample range 7.9%–10.7%n=900
  11. 10.8%
    Waypoint AirGrok 4 · CrewAIExample range 9.8%–11.8%n=900
  12. 14.1%
    Task PilotMeta · CrewAIExample range 12.7%–12.0%n=900

How much more the booked flight cost than the cheapest valid option after bags, seats, taxes, and required fees.

Complete agents

Overall score

Higher is betterExample score out of 100
RoutekeeperGPT-5 · OpenAI Agents SDK
Overall
73.2
Rank
1/12
Range
71.5–74.9
Runs
900
  1. 73.2
    RoutekeeperGPT-5 · OpenAI Agents SDKExample range 71.5–74.9n=900
  2. 72.5
    Aero GuideClaude Opus 4.1 · LangGraphExample range 70.7–74.3n=900
  3. 70.8
    Flight LensGemini 2.5 Pro · Google ADKExample range 68.9–72.7n=900
  4. 70.0
    Compass AgentOpenAI · Agents SDKExample range 68.2–71.7n=900
  5. 68.4
    Scout AIAnthropic · LangGraphExample range 66.7–70.1n=900
  6. 66.1
    Open SkiesMistral Large 2 · PydanticAIExample range 64.2–68.0n=900
  7. 65.8
    Relay ConciergeGoogle · Google ADKExample range 64.2–67.5n=900
  8. 65.7
    Nova AssistDeepSeek · LangChainExample range 64.1–67.3n=900
  9. 64.2
    Waypoint ProxAI · Native toolsExample range 62.5–65.8n=900
  10. 62.4
    Waypoint AirGrok 4 · CrewAIExample range 60.4–64.4n=900
  11. 60.3
    Orbit PlannerMistral · PydanticAIExample range 58.8–61.8n=900
  12. 55.8
    Task PilotMeta · CrewAIExample range 54.4–57.1n=900

A simple score combining task success, following requirements, recovery, cost, and trust.

Complete agents

Rebooked after a disruption

Higher is betterRuns where something went wrong
Aero GuideClaude Opus 4.1 · LangGraph
Recovery
71%
Rank
1/12
Range
68%–74%
Runs
900
  1. 71%
    Aero GuideClaude Opus 4.1 · LangGraphExample range 68%–74%n=900
  2. 67%
    RoutekeeperGPT-5 · OpenAI Agents SDKExample range 64%–70%n=900
  3. 67%
    Scout AIAnthropic · LangGraphExample range 65%–69%n=900
  4. 64%
    Flight LensGemini 2.5 Pro · Google ADKExample range 61%–67%n=900
  5. 64%
    Compass AgentOpenAI · Agents SDKExample range 62%–65%n=900
  6. 63%
    Waypoint ProxAI · Native toolsExample range 61%–64%n=900
  7. 61%
    Open SkiesMistral Large 2 · PydanticAIExample range 58%–64%n=900
  8. 60%
    Nova AssistDeepSeek · LangChainExample range 58%–61%n=900
  9. 59%
    Relay ConciergeGoogle · Google ADKExample range 58%–61%n=900
  10. 56%
    Waypoint AirGrok 4 · CrewAIExample range 53%–59%n=900
  11. 55%
    Orbit PlannerMistral · PydanticAIExample range 54%–57%n=900
  12. 49%
    Task PilotMeta · CrewAIExample range 48%–51%n=900

How often the agent still finished after a price change, tool error, rejection, or timeout.

Complete agents

Followed approval rules

Higher is betterExample score out of 100
Aero GuideClaude Opus 4.1 · LangGraph
Trust
88%
Rank
1/12
Range
86%–90%
Runs
900
  1. 88%
    Aero GuideClaude Opus 4.1 · LangGraphExample range 86%–90%n=900
  2. 84%
    RoutekeeperGPT-5 · OpenAI Agents SDKExample range 82%–86%n=900
  3. 84%
    Scout AIAnthropic · LangGraphExample range 82%–86%n=900
  4. 82%
    Flight LensGemini 2.5 Pro · Google ADKExample range 80%–84%n=900
  5. 81%
    Compass AgentOpenAI · Agents SDKExample range 79%–83%n=900
  6. 80%
    Waypoint ProxAI · Native toolsExample range 78%–82%n=900
  7. 77%
    Relay ConciergeGoogle · Google ADKExample range 75%–79%n=900
  8. 77%
    Open SkiesMistral Large 2 · PydanticAIExample range 75%–79%n=900
  9. 77%
    Nova AssistDeepSeek · LangChainExample range 75%–78%n=900
  10. 72%
    Waypoint AirGrok 4 · CrewAIExample range 70%–74%n=900
  11. 71%
    Orbit PlannerMistral · PydanticAIExample range 69%–73%n=900
  12. 65%
    Task PilotMeta · CrewAIExample range 64%–67%n=900

Whether the agent asked before important actions, showed its proof, and was honest about uncertainty.

Speed and cost

Time vs. cost

Best area is lower-leftCircle size · Correct flights booked
Time vs. costBest area is lower-left. Flight Lens is closest to that area in this example, at 7m 16s fastest booking and $0.88 cheapest booking. Circle size shows correct flights booked.Best area6m 27s8m 19s10m 10s12m 02s13m 53s$0.59$0.94$1.29$1.64$2.00Fastest booking · Lower is betterCheapest booking · Lower is betterRoutekeeper: Fastest booking 7m 48s; Cheapest booking $1.14; Correct flights booked 71%1Aero Guide: Fastest booking 8m 32s; Cheapest booking $1.32; Correct flights booked 69%2Flight Lens: Fastest booking 7m 16s; Cheapest booking $0.88; Correct flights booked 68%3Open Skies: Fastest booking 9m 08s; Cheapest booking $0.74; Correct flights booked 63%4Waypoint Air: Fastest booking 10m 01s; Cheapest booking $1.08; Correct flights booked 59%5Compass Agent: Fastest booking 8m 47s; Cheapest booking $1.28; Correct flights booked 68%6Scout AI: Fastest booking 9m 59s; Cheapest booking $1.54; Correct flights booked 65%7Relay Concierge: Fastest booking 8m 50s; Cheapest booking $1.07; Correct flights booked 63%8Orbit Planner: Fastest booking 11m 30s; Cheapest booking $0.93; Correct flights booked 57%9Task Pilot: Fastest booking 13m 04s; Cheapest booking $1.41; Correct flights booked 52%10Nova Assist: Fastest booking 10m 32s; Cheapest booking $1.54; Correct flights booked 64%11Waypoint Pro: Fastest booking 11m 54s; Cheapest booking $1.84; Correct flights booked 61%12
  1. 1Routekeeper7m 48s × $1.14 · Booked 71%
  2. 2Aero Guide8m 32s × $1.32 · Booked 69%
  3. 3Flight Lens7m 16s × $0.88 · Booked 68%
  4. 4Open Skies9m 08s × $0.74 · Booked 63%
  5. 5Waypoint Air10m 01s × $1.08 · Booked 59%
  6. 6Compass Agent8m 47s × $1.28 · Booked 68%
  7. 7Scout AI9m 59s × $1.54 · Booked 65%
  8. 8Relay Concierge8m 50s × $1.07 · Booked 63%
  9. 9Orbit Planner11m 30s × $0.93 · Booked 57%
  10. 10Task Pilot13m 04s × $1.41 · Booked 52%
  11. 11Nova Assist10m 32s × $1.54 · Booked 64%
  12. 12Waypoint Pro11m 54s × $1.84 · Booked 61%

Best area is lower-left. Flight Lens is closest to that area in this example, at 7m 16s fastest booking and $0.88 cheapest booking. Circle size shows correct flights booked.

Agents closer to the lower-left finished faster and cost less. Bigger circles booked more correct flights.

About this dataMade-up examples for this preview—not real test resultsRangeGrouped by flight taskProof requiredReceipt or other outside proofDataMade-up examples for this preview
02

Models only

300 tasks × 3 tries

Each model sees the same task and options. Tools are turned off, so the model cannot take actions.

Same test for everyoneSame task, options, and scoring. Tools are off.

Main result

Overall reasoning

Higher is betterTools off · example score out of 100
GPT-5OpenAI · tools disabled
Overall
86.7
Rank
1/12
Range
85.3–88.1
Runs
900
  1. 86.7
    GPT-5OpenAI · tools disabledExample range 85.3–88.1n=900
  2. 85.9
    Claude Opus 4.1Anthropic · tools disabledExample range 84.4–87.4n=900
  3. 83.5
    GPT-4.1OpenAI · tools disabledExample range 81.4–85.5n=900
  4. 82.4
    Gemini 2.5 ProGoogle · tools disabledExample range 80.8–84.0n=900
  5. 81.8
    Claude 3.7 SonnetAnthropic · tools disabledExample range 79.8–83.8n=900
  6. 79.2
    DeepSeek V3DeepSeek · tools disabledExample range 77.2–81.2n=900
  7. 77.6
    Grok 3xAI · tools disabledExample range 75.6–79.5n=900
  8. 77.5
    Gemini 2.0 FlashGoogle · tools disabledExample range 75.5–79.4n=900
  9. 75.8
    Mistral Large 2Mistral AI · tools disabledExample range 74.2–77.4n=900
  10. 73.2
    Grok 4xAI · tools disabledExample range 71.5–74.9n=900
  11. 70.0
    Mistral LargeMistral · tools disabledExample range 68.3–71.8n=900
  12. 66.5
    Llama 4 MaverickMeta · tools disabledExample range 64.9–68.2n=900

How well the model understood the task and made a choice. The model could not use tools or take actions.

Models only

Remembered requirements

Higher is betterSame task information for every model
Claude Opus 4.1Anthropic · tools disabled
Requirements
96%
Rank
1/12
Range
94%–98%
Runs
900
  1. 96%
    Claude Opus 4.1Anthropic · tools disabledExample range 94%–98%n=900
  2. 94%
    GPT-5OpenAI · tools disabledExample range 92%–96%n=900
  3. 92%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 90%–94%n=900
  4. 91%
    Gemini 2.5 ProGoogle · tools disabledExample range 89%–93%n=900
  5. 91%
    GPT-4.1OpenAI · tools disabledExample range 88%–93%n=900
  6. 88%
    Grok 3xAI · tools disabledExample range 85%–90%n=900
  7. 87%
    DeepSeek V3DeepSeek · tools disabledExample range 84%–89%n=900
  8. 86%
    Gemini 2.0 FlashGoogle · tools disabledExample range 84%–88%n=900
  9. 86%
    Mistral Large 2Mistral AI · tools disabledExample range 84%–88%n=900
  10. 82%
    Grok 4xAI · tools disabledExample range 80%–84%n=900
  11. 80%
    Mistral LargeMistral · tools disabledExample range 78%–82%n=900
  12. 75%
    Llama 4 MaverickMeta · tools disabledExample range 73%–77%n=900

How often the model remembered the user’s must-haves when making a choice.

Models only

Choice quality

Higher is betterSame options and scoring
GPT-5OpenAI · tools disabled
Choice
88%
Rank
1/12
Range
86%–90%
Runs
900
  1. 88%
    GPT-5OpenAI · tools disabledExample range 86%–90%n=900
  2. 87%
    Claude Opus 4.1Anthropic · tools disabledExample range 85%–89%n=900
  3. 85%
    GPT-4.1OpenAI · tools disabledExample range 83%–87%n=900
  4. 84%
    Gemini 2.5 ProGoogle · tools disabledExample range 82%–86%n=900
  5. 83%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 81%–85%n=900
  6. 81%
    DeepSeek V3DeepSeek · tools disabledExample range 78%–83%n=900
  7. 79%
    Gemini 2.0 FlashGoogle · tools disabledExample range 77%–81%n=900
  8. 79%
    Grok 3xAI · tools disabledExample range 77%–81%n=900
  9. 76%
    Mistral Large 2Mistral AI · tools disabledExample range 74%–78%n=900
  10. 75%
    Grok 4xAI · tools disabledExample range 73%–77%n=900
  11. 70%
    Mistral LargeMistral · tools disabledExample range 68%–72%n=900
  12. 68%
    Llama 4 MaverickMeta · tools disabledExample range 67%–70%n=900

How good the model’s choice was when every model saw the same options.

Models only

Action plan works

Higher is betterPlans reviewed without running them
GPT-5OpenAI · tools disabled
Plan
84%
Rank
1/12
Range
82%–86%
Runs
900
  1. 84%
    GPT-5OpenAI · tools disabledExample range 82%–86%n=900
  2. 82%
    Claude Opus 4.1Anthropic · tools disabledExample range 80%–84%n=900
  3. 81%
    GPT-4.1OpenAI · tools disabledExample range 79%–83%n=900
  4. 80%
    Gemini 2.5 ProGoogle · tools disabledExample range 78%–82%n=900
  5. 78%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 76%–80%n=900
  6. 77%
    DeepSeek V3DeepSeek · tools disabledExample range 75%–78%n=900
  7. 75%
    Gemini 2.0 FlashGoogle · tools disabledExample range 73%–77%n=900
  8. 74%
    Grok 3xAI · tools disabledExample range 72%–75%n=900
  9. 73%
    Mistral Large 2Mistral AI · tools disabledExample range 71%–76%n=900
  10. 71%
    Grok 4xAI · tools disabledExample range 68%–74%n=900
  11. 67%
    Mistral LargeMistral · tools disabledExample range 66%–69%n=900
  12. 64%
    Llama 4 MaverickMeta · tools disabledExample range 63%–66%n=900

How often the proposed steps were possible, in the right order, and safe to follow.

Models only

Honest about uncertainty

Higher is betterSame scoring for every model
Claude Opus 4.1Anthropic · tools disabled
Honesty
89%
Rank
1/12
Range
87%–91%
Runs
900
  1. 89%
    Claude Opus 4.1Anthropic · tools disabledExample range 87%–91%n=900
  2. 85%
    GPT-5OpenAI · tools disabledExample range 83%–87%n=900
  3. 85%
    Claude 3.7 SonnetAnthropic · tools disabledExample range 83%–87%n=900
  4. 83%
    Gemini 2.5 ProGoogle · tools disabledExample range 81%–85%n=900
  5. 82%
    GPT-4.1OpenAI · tools disabledExample range 80%–84%n=900
  6. 81%
    Grok 3xAI · tools disabledExample range 79%–83%n=900
  7. 78%
    Gemini 2.0 FlashGoogle · tools disabledExample range 76%–80%n=900
  8. 78%
    Mistral Large 2Mistral AI · tools disabledExample range 76%–80%n=900
  9. 78%
    DeepSeek V3DeepSeek · tools disabledExample range 76%–79%n=900
  10. 72%
    Mistral LargeMistral · tools disabledExample range 70%–74%n=900
  11. 72%
    Grok 4xAI · tools disabledExample range 70%–74%n=900
  12. 65%
    Llama 4 MaverickMeta · tools disabledExample range 64%–67%n=900

Whether the model used the available proof and avoided claiming success when it was unsure.

Speed and cost

Good choice vs. workable plan

Best area is upper-rightCircle size · Honest about uncertainty
Good choice vs. workable planBest area is upper-right. GPT-5 is closest to that area in this example, at 84% action plan works and 88% choice quality. Circle size shows honest about uncertainty.Best area0%25%50%75%100%0%25%50%75%100%Action plan works · Higher is betterChoice quality · Higher is betterGPT-5: Action plan works 84%; Choice quality 88%; Honest about uncertainty 85%1Claude Opus 4.1: Action plan works 82%; Choice quality 87%; Honest about uncertainty 89%2Gemini 2.5 Pro: Action plan works 80%; Choice quality 84%; Honest about uncertainty 83%3Mistral Large 2: Action plan works 73%; Choice quality 76%; Honest about uncertainty 78%4Grok 4: Action plan works 71%; Choice quality 75%; Honest about uncertainty 72%5GPT-4.1: Action plan works 81%; Choice quality 85%; Honest about uncertainty 82%6Claude 3.7 Sonnet: Action plan works 78%; Choice quality 83%; Honest about uncertainty 85%7Gemini 2.0 Flash: Action plan works 75%; Choice quality 79%; Honest about uncertainty 78%8Mistral Large: Action plan works 67%; Choice quality 70%; Honest about uncertainty 72%9Llama 4 Maverick: Action plan works 64%; Choice quality 68%; Honest about uncertainty 65%10DeepSeek V3: Action plan works 77%; Choice quality 81%; Honest about uncertainty 78%11Grok 3: Action plan works 74%; Choice quality 79%; Honest about uncertainty 81%12
  1. 1GPT-584% × 88% · Honesty 85%
  2. 2Claude Opus 4.182% × 87% · Honesty 89%
  3. 3Gemini 2.5 Pro80% × 84% · Honesty 83%
  4. 4Mistral Large 273% × 76% · Honesty 78%
  5. 5Grok 471% × 75% · Honesty 72%
  6. 6GPT-4.181% × 85% · Honesty 82%
  7. 7Claude 3.7 Sonnet78% × 83% · Honesty 85%
  8. 8Gemini 2.0 Flash75% × 79% · Honesty 78%
  9. 9Mistral Large67% × 70% · Honesty 72%
  10. 10Llama 4 Maverick64% × 68% · Honesty 65%
  11. 11DeepSeek V377% × 81% · Honesty 78%
  12. 12Grok 374% × 79% · Honesty 81%

Best area is upper-right. GPT-5 is closest to that area in this example, at 84% action plan works and 88% choice quality. Circle size shows honest about uncertainty.

Models in the upper-right make better choices and produce plans that are more likely to work. Bigger circles mean the model was more honest about uncertainty.

About this dataMade-up examples for this preview—not real test resultsActionsTurned offOptionsSame options for every modelDataMade-up examples for this preview
03

Agent setup

300 tasks × 7 problem types

We keep the model and tools the same and change only the workflow that runs them.

Same test for everyoneSame model, tools, task, and problems.

Main result

Agent setup score

Higher is betterSame model and tools · example score out of 100
LangGraphLangChain · frozen model and adapters
Overall
83.8
Rank
1/12
Range
82.3–85.3
Runs
900
  1. 83.8
    LangGraphLangChain · frozen model and adaptersExample range 82.3–85.3n=900
  2. 82.7
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 81.1–84.3n=900
  3. 80.5
    Agents SDKOpenAI · Agents SDKExample range 78.5–82.6n=900
  4. 78.9
    Google ADKGoogle · frozen model and adaptersExample range 77.2–80.6n=900
  5. 78.6
    LangGraphLangChain · LangGraphExample range 76.6–80.6n=900
  6. 77.4
    PydanticAIPydantic · frozen model and adaptersExample range 75.7–79.1n=900
  7. 77.1
    CrewAICrewAI · CrewAIExample range 75.2–79.1n=900
  8. 75.2
    Native toolsAnthropic · Native toolsExample range 73.3–77.1n=900
  9. 74.0
    PydanticAIPydantic · PydanticAIExample range 72.1–75.8n=900
  10. 71.6
    Google ADKGoogle · Google ADKExample range 69.8–73.4n=900
  11. 70.6
    BrowserbaseBrowserbase · BrowserbaseExample range 68.8–72.3n=900
  12. 68.2
    Agents SDKOpenAI · Agents SDKExample range 66.5–69.9n=900

How well the workflow ran the task when the model and tools stayed the same.

Agent setup

Ran the plan correctly

Higher is betterSame approved plan
OpenAI Agents SDKOpenAI · frozen model and adapters
Execution
84%
Rank
1/12
Range
82%–86%
Runs
900
  1. 84%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 82%–86%n=900
  2. 81%
    LangGraphLangChain · frozen model and adaptersExample range 79%–83%n=900
  3. 80%
    LangGraphLangChain · LangGraphExample range 78%–82%n=900
  4. 79%
    Google ADKGoogle · frozen model and adaptersExample range 77%–81%n=900
  5. 78%
    Agents SDKOpenAI · Agents SDKExample range 76%–80%n=900
  6. 77%
    Native toolsAnthropic · Native toolsExample range 75%–78%n=900
  7. 76%
    PydanticAIPydantic · frozen model and adaptersExample range 74%–78%n=900
  8. 74%
    CrewAICrewAI · CrewAIExample range 72%–76%n=900
  9. 74%
    PydanticAIPydantic · PydanticAIExample range 72%–76%n=900
  10. 71%
    BrowserbaseBrowserbase · BrowserbaseExample range 69%–72%n=900
  11. 70%
    Google ADKGoogle · Google ADKExample range 68%–72%n=900
  12. 67%
    Agents SDKOpenAI · Agents SDKExample range 65%–68%n=900

How often the setup followed the approved steps in the right order without losing progress.

Agent setup

Recovered from an error

Higher is betterRuns where something went wrong
LangGraphLangChain · frozen model and adapters
Recovery
80%
Rank
1/12
Range
77%–83%
Runs
900
  1. 80%
    LangGraphLangChain · frozen model and adaptersExample range 77%–83%n=900
  2. 77%
    Agents SDKOpenAI · Agents SDKExample range 75%–79%n=900
  3. 74%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 71%–77%n=900
  4. 73%
    CrewAICrewAI · CrewAIExample range 72%–75%n=900
  5. 73%
    PydanticAIPydantic · frozen model and adaptersExample range 70%–76%n=900
  6. 71%
    Google ADKGoogle · frozen model and adaptersExample range 68%–74%n=900
  7. 70%
    LangGraphLangChain · LangGraphExample range 68%–72%n=900
  8. 67%
    Google ADKGoogle · Google ADKExample range 66%–69%n=900
  9. 67%
    Native toolsAnthropic · Native toolsExample range 65%–68%n=900
  10. 66%
    PydanticAIPydantic · PydanticAIExample range 64%–68%n=900
  11. 64%
    Agents SDKOpenAI · Agents SDKExample range 62%–65%n=900
  12. 63%
    BrowserbaseBrowserbase · BrowserbaseExample range 61%–64%n=900

How often the setup recovered from a timeout, outdated information, price change, or rejection.

Agent setup

Remembered progress

Higher is betterProgress record checked
LangGraphLangChain · frozen model and adapters
Memory
93%
Rank
1/12
Range
91%–95%
Runs
900
  1. 93%
    LangGraphLangChain · frozen model and adaptersExample range 91%–95%n=900
  2. 90%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 88%–92%n=900
  3. 90%
    Agents SDKOpenAI · Agents SDKExample range 88%–92%n=900
  4. 89%
    PydanticAIPydantic · frozen model and adaptersExample range 87%–91%n=900
  5. 87%
    Google ADKGoogle · frozen model and adaptersExample range 85%–89%n=900
  6. 86%
    CrewAICrewAI · CrewAIExample range 84%–89%n=900
  7. 86%
    LangGraphLangChain · LangGraphExample range 84%–88%n=900
  8. 83%
    Google ADKGoogle · Google ADKExample range 81%–85%n=900
  9. 83%
    Native toolsAnthropic · Native toolsExample range 80%–85%n=900
  10. 82%
    PydanticAIPydantic · PydanticAIExample range 80%–84%n=900
  11. 80%
    Agents SDKOpenAI · Agents SDKExample range 78%–82%n=900
  12. 79%
    BrowserbaseBrowserbase · BrowserbaseExample range 77%–81%n=900

Whether the setup remembered requirements, approvals, evidence, and completed steps.

Agent setup

Waited for approval

Higher is betterApproval checks
OpenAI Agents SDKOpenAI · frozen model and adapters
Approval
96%
Rank
1/12
Range
94%–98%
Runs
900
  1. 96%
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range 94%–98%n=900
  2. 94%
    LangGraphLangChain · frozen model and adaptersExample range 92%–96%n=900
  3. 93%
    PydanticAIPydantic · frozen model and adaptersExample range 91%–95%n=900
  4. 92%
    Google ADKGoogle · frozen model and adaptersExample range 90%–94%n=900
  5. 92%
    LangGraphLangChain · LangGraphExample range 90%–94%n=900
  6. 91%
    Agents SDKOpenAI · Agents SDKExample range 88%–93%n=900
  7. 89%
    Native toolsAnthropic · Native toolsExample range 86%–91%n=900
  8. 87%
    CrewAICrewAI · CrewAIExample range 85%–90%n=900
  9. 87%
    Google ADKGoogle · Google ADKExample range 85%–89%n=900
  10. 87%
    PydanticAIPydantic · PydanticAIExample range 85%–89%n=900
  11. 84%
    Agents SDKOpenAI · Agents SDKExample range 82%–86%n=900
  12. 84%
    BrowserbaseBrowserbase · BrowserbaseExample range 82%–86%n=900

How often the setup asked the user before an important or hard-to-reverse action.

Cost

Extra setup cost

Lower is betterExample USD above the baseline
PydanticAIPydantic · frozen model and adapters
Extra cost
$0.11
Rank
1/12
Range
$0.08–$0.14
Runs
900
  1. $0.11
    PydanticAIPydantic · frozen model and adaptersExample range $0.08–$0.14n=900
  2. $0.14
    Google ADKGoogle · Google ADKExample range $0.11–$0.17n=900
  3. $0.14
    Google ADKGoogle · frozen model and adaptersExample range $0.11–$0.17n=900
  4. $0.16
    Agents SDKOpenAI · Agents SDKExample range $0.13–$0.19n=900
  5. $0.17
    PydanticAIPydantic · PydanticAIExample range $0.14–$0.20n=900
  6. $0.18
    OpenAI Agents SDKOpenAI · frozen model and adaptersExample range $0.16–$0.20n=900
  7. $0.20
    BrowserbaseBrowserbase · BrowserbaseExample range $0.17–$0.23n=900
  8. $0.21
    LangGraphLangChain · frozen model and adaptersExample range $0.19–$0.23n=900
  9. $0.21
    LangGraphLangChain · LangGraphExample range $0.18–$0.24n=900
  10. $0.24
    Agents SDKOpenAI · Agents SDKExample range $0.21–$0.27n=900
  11. $0.24
    Native toolsAnthropic · Native toolsExample range $0.21–$0.27n=900
  12. $0.27
    CrewAICrewAI · CrewAIExample range $0.24–$0.30n=900

The extra cost added by the agent setup for each completed task.

Speed and cost

Memory vs. extra cost

Best area is upper-leftCircle size · Waited for approval
Memory vs. extra costBest area is upper-left. PydanticAI is closest to that area in this example, at $0.11 extra setup cost and 89% remembered progress. Circle size shows waited for approval.Best area$0.09$0.14$0.19$0.24$0.300%25%50%75%100%Extra setup cost · Lower is betterRemembered progress · Higher is betterLangGraph: Extra setup cost $0.21; Remembered progress 93%; Waited for approval 94%1OpenAI Agents SDK: Extra setup cost $0.18; Remembered progress 90%; Waited for approval 96%2Google ADK: Extra setup cost $0.14; Remembered progress 87%; Waited for approval 92%3PydanticAI: Extra setup cost $0.11; Remembered progress 89%; Waited for approval 93%4Agents SDK: Extra setup cost $0.24; Remembered progress 90%; Waited for approval 91%5LangGraph: Extra setup cost $0.21; Remembered progress 86%; Waited for approval 92%6PydanticAI: Extra setup cost $0.17; Remembered progress 82%; Waited for approval 87%7Google ADK: Extra setup cost $0.14; Remembered progress 83%; Waited for approval 87%8CrewAI: Extra setup cost $0.27; Remembered progress 86%; Waited for approval 87%9Native tools: Extra setup cost $0.24; Remembered progress 83%; Waited for approval 89%10Browserbase: Extra setup cost $0.20; Remembered progress 79%; Waited for approval 84%11Agents SDK: Extra setup cost $0.16; Remembered progress 80%; Waited for approval 84%12
  1. 1LangGraph$0.21 × 93% · Approval 94%
  2. 2OpenAI Agents SDK$0.18 × 90% · Approval 96%
  3. 3Google ADK$0.14 × 87% · Approval 92%
  4. 4PydanticAI$0.11 × 89% · Approval 93%
  5. 5Agents SDK$0.24 × 90% · Approval 91%
  6. 6LangGraph$0.21 × 86% · Approval 92%
  7. 7PydanticAI$0.17 × 82% · Approval 87%
  8. 8Google ADK$0.14 × 83% · Approval 87%
  9. 9CrewAI$0.27 × 86% · Approval 87%
  10. 10Native tools$0.24 × 83% · Approval 89%
  11. 11Browserbase$0.20 × 79% · Approval 84%
  12. 12Agents SDK$0.16 × 80% · Approval 84%

Best area is upper-left. PydanticAI is closest to that area in this example, at $0.11 extra setup cost and 89% remembered progress. Circle size shows waited for approval.

Setups in the upper-left remember more while adding less cost. Bigger circles mean they waited for approval more reliably.

About this dataMade-up examples for this preview—not real test resultsModelSame model for every setupProblemsTimeout, old information, rejectionDataMade-up examples for this preview
04

Connected tools

300 requests × 3 tries

We send the same requests to each tool and check whether the action worked, the information was current, and proof came back.

Same test for everyoneSame model, agent setup, request, retry rules, and scoring.

Main result

Tool reliability

Higher is betterSame model and setup · example score out of 100
Amadeus flight connectorSearch, pricing, and order creation
Overall
91.6
Rank
1/12
Range
90.3–92.9
Runs
900
  1. 91.6
    Amadeus flight connectorSearch, pricing, and order creationExample range 90.3–92.9n=900
  2. 88.7
    Skyscanner flight connectorMetasearch and airline handoffExample range 87.3–90.1n=900
  3. 88.3
    Chrome connectorGoogle · ChromeExample range 86.1–90.6n=900
  4. 87.8
    Google Flights connectorMetasearch and fare comparisonExample range 86.3–89.3n=900
  5. 85.9
    Duffel flight connectorSearch, pricing, and order creationExample range 84.4–87.4n=900
  6. 84.6
    Playwright connectorMicrosoft · PlaywrightExample range 82.5–86.7n=900
  7. 84.1
    Stripe connectorStripe · StripeExample range 82.0–86.2n=900
  8. 82.8
    Gmail connectorGoogle · GmailExample range 80.8–84.9n=900
  9. 82.3
    KAYAK flight connectorMetasearch and airline handoffExample range 80.7–83.9n=900
  10. 80.4
    Zendesk connectorZendesk · ZendeskExample range 78.3–82.4n=900
  11. 80.1
    Twilio connectorTwilio · TwilioExample range 78.1–82.1n=900
  12. 75.6
    Mapbox connectorMapbox · MapboxExample range 73.8–77.5n=900

How reliably the connected tool completed actions, returned current information, and provided proof.

Connected tools

Action worked

Higher is betterSame requests for every tool
Amadeus flight connectorSearch, pricing, and order creation
Worked
92%
Rank
1/12
Range
90%–94%
Runs
900
  1. 92%
    Amadeus flight connectorSearch, pricing, and order creationExample range 90%–94%n=900
  2. 89%
    Chrome connectorGoogle · ChromeExample range 87%–91%n=900
  3. 88%
    Skyscanner flight connectorMetasearch and airline handoffExample range 86%–90%n=900
  4. 87%
    Google Flights connectorMetasearch and fare comparisonExample range 85%–89%n=900
  5. 86%
    Duffel flight connectorSearch, pricing, and order creationExample range 84%–88%n=900
  6. 85%
    Stripe connectorStripe · StripeExample range 82%–87%n=900
  7. 84%
    Playwright connectorMicrosoft · PlaywrightExample range 82%–86%n=900
  8. 82%
    Gmail connectorGoogle · GmailExample range 80%–84%n=900
  9. 82%
    KAYAK flight connectorMetasearch and airline handoffExample range 80%–84%n=900
  10. 80%
    Twilio connectorTwilio · TwilioExample range 78%–82%n=900
  11. 80%
    Zendesk connectorZendesk · ZendeskExample range 78%–82%n=900
  12. 75%
    Mapbox connectorMapbox · MapboxExample range 73%–77%n=900

How often the tool completed the requested action and returned proof.

Connected tools

Information was current

Higher is betterChecked against the same source
Google Flights connectorMetasearch and fare comparison
Current info
97%
Rank
1/12
Range
95%–99%
Runs
900
  1. 97%
    Google Flights connectorMetasearch and fare comparisonExample range 95%–99%n=900
  2. 95%
    Amadeus flight connectorSearch, pricing, and order creationExample range 93%–97%n=900
  3. 94%
    Skyscanner flight connectorMetasearch and airline handoffExample range 92%–96%n=900
  4. 92%
    Gmail connectorGoogle · GmailExample range 90%–94%n=900
  5. 92%
    Chrome connectorGoogle · ChromeExample range 89%–94%n=900
  6. 91%
    Duffel flight connectorSearch, pricing, and order creationExample range 89%–93%n=900
  7. 90%
    Playwright connectorMicrosoft · PlaywrightExample range 88%–92%n=900
  8. 89%
    KAYAK flight connectorMetasearch and airline handoffExample range 87%–91%n=900
  9. 88%
    Stripe connectorStripe · StripeExample range 85%–90%n=900
  10. 86%
    Zendesk connectorZendesk · ZendeskExample range 84%–88%n=900
  11. 85%
    Twilio connectorTwilio · TwilioExample range 83%–87%n=900
  12. 82%
    Mapbox connectorMapbox · MapboxExample range 80%–84%n=900

How often the tool returned the correct price, availability, or policy at that moment.

Connected tools

Retry worked

Higher is betterRuns with a temporary error
Amadeus flight connectorSearch, pricing, and order creation
Retry
81%
Rank
1/12
Range
79%–84%
Runs
900
  1. 81%
    Amadeus flight connectorSearch, pricing, and order creationExample range 79%–84%n=900
  2. 78%
    Skyscanner flight connectorMetasearch and airline handoffExample range 75%–81%n=900
  3. 78%
    Chrome connectorGoogle · ChromeExample range 76%–80%n=900
  4. 76%
    Duffel flight connectorSearch, pricing, and order creationExample range 73%–79%n=900
  5. 74%
    Google Flights connectorMetasearch and fare comparisonExample range 71%–77%n=900
  6. 74%
    Playwright connectorMicrosoft · PlaywrightExample range 72%–76%n=900
  7. 74%
    Stripe connectorStripe · StripeExample range 72%–75%n=900
  8. 70%
    Twilio connectorTwilio · TwilioExample range 68%–72%n=900
  9. 70%
    Zendesk connectorZendesk · ZendeskExample range 68%–71%n=900
  10. 69%
    Gmail connectorGoogle · GmailExample range 67%–71%n=900
  11. 69%
    KAYAK flight connectorMetasearch and airline handoffExample range 66%–72%n=900
  12. 62%
    Mapbox connectorMapbox · MapboxExample range 61%–64%n=900

How often the tool recovered after a temporary error.

Connected tools

Proof returned

Higher is betterReceipts and status details checked
Amadeus flight connectorSearch, pricing, and order creation
Proof
96%
Rank
1/12
Range
94%–98%
Runs
900
  1. 96%
    Amadeus flight connectorSearch, pricing, and order creationExample range 94%–98%n=900
  2. 93%
    Chrome connectorGoogle · ChromeExample range 90%–95%n=900
  3. 92%
    Duffel flight connectorSearch, pricing, and order creationExample range 90%–94%n=900
  4. 90%
    Skyscanner flight connectorMetasearch and airline handoffExample range 88%–92%n=900
  5. 89%
    Stripe connectorStripe · StripeExample range 86%–91%n=900
  6. 88%
    Google Flights connectorMetasearch and fare comparisonExample range 86%–90%n=900
  7. 86%
    Twilio connectorTwilio · TwilioExample range 84%–88%n=900
  8. 86%
    KAYAK flight connectorMetasearch and airline handoffExample range 84%–88%n=900
  9. 86%
    Playwright connectorMicrosoft · PlaywrightExample range 84%–88%n=900
  10. 83%
    Gmail connectorGoogle · GmailExample range 81%–85%n=900
  11. 82%
    Zendesk connectorZendesk · ZendeskExample range 80%–84%n=900
  12. 79%
    Mapbox connectorMapbox · MapboxExample range 77%–81%n=900

How often the tool returned the details needed to prove what happened.

Speed and cost

Current information vs. proof

Best area is upper-rightCircle size · Retry worked
Current information vs. proofBest area is upper-right. Amadeus flight connector is closest to that area in this example, at 95% information was current and 96% proof returned. Circle size shows retry worked.Best area0%25%50%75%100%0%25%50%75%100%Information was current · Higher is betterProof returned · Higher is betterAmadeus flight connector: Information was current 95%; Proof returned 96%; Retry worked 81%1Skyscanner flight connector: Information was current 94%; Proof returned 90%; Retry worked 78%2Google Flights connector: Information was current 97%; Proof returned 88%; Retry worked 74%3Duffel flight connector: Information was current 91%; Proof returned 92%; Retry worked 76%4KAYAK flight connector: Information was current 89%; Proof returned 86%; Retry worked 69%5Chrome connector: Information was current 92%; Proof returned 93%; Retry worked 78%6Playwright connector: Information was current 90%; Proof returned 86%; Retry worked 74%7Gmail connector: Information was current 92%; Proof returned 83%; Retry worked 69%8Twilio connector: Information was current 85%; Proof returned 86%; Retry worked 70%9Mapbox connector: Information was current 82%; Proof returned 79%; Retry worked 62%10Stripe connector: Information was current 88%; Proof returned 89%; Retry worked 74%11Zendesk connector: Information was current 86%; Proof returned 82%; Retry worked 70%12
  1. 1Amadeus flight connector95% × 96% · Retry 81%
  2. 2Skyscanner flight connector94% × 90% · Retry 78%
  3. 3Google Flights connector97% × 88% · Retry 74%
  4. 4Duffel flight connector91% × 92% · Retry 76%
  5. 5KAYAK flight connector89% × 86% · Retry 69%
  6. 6Chrome connector92% × 93% · Retry 78%
  7. 7Playwright connector90% × 86% · Retry 74%
  8. 8Gmail connector92% × 83% · Retry 69%
  9. 9Twilio connector85% × 86% · Retry 70%
  10. 10Mapbox connector82% × 79% · Retry 62%
  11. 11Stripe connector88% × 89% · Retry 74%
  12. 12Zendesk connector86% × 82% · Retry 70%

Best area is upper-right. Amadeus flight connector is closest to that area in this example, at 95% information was current and 96% proof returned. Circle size shows retry worked.

Tools in the upper-right return more current information and better proof. Bigger circles mean retries worked more often.

About this dataMade-up examples for this preview—not real test resultsRequestsSame approved requestsStarting pointSame prices, availability, and rulesDataMade-up examples for this preview

Detailed comparison

See what changes the result.

Compare the complete agent, the model alone, the agent setup, and the connected tools.

Complete agent

Which agent actually finishes the job?

We test the model, agent setup, and tools together. It only counts when we can verify the result.

Same test for everyoneEvery agent gets the same person, task, rules, and starting information.

Ranking

Overall score

Sample output for interface demonstration
01
Routekeeper
GPT-5 · OpenAI Agents SDK
73.2
02
Aero Guide
Claude Opus 4.1 · LangGraph
72.5
03
Flight Lens
Gemini 2.5 Pro · Google ADK
70.8
04
Open Skies
Mistral Large 2 · PydanticAI
66.1
05
Waypoint Air
Grok 4 · CrewAI
62.4

Higher is better. Small score differences may not matter once real test results replace these examples.

RankComplete agentsOverallCompletedRecoveryTrustCostTime
1
RoutekeeperGPT-5 · OpenAI Agents SDK
73.271.067.084.0$1.147m 48s
2
Aero GuideClaude Opus 4.1 · LangGraph
72.569.071.088.0$1.328m 32s
3
Flight LensGemini 2.5 Pro · Google ADK
70.868.064.082.0$0.887m 16s
4
Open SkiesMistral Large 2 · PydanticAI
66.163.061.077.0$0.749m 08s
5
Waypoint AirGrok 4 · CrewAI
62.459.056.072.0$1.0810m 01s
What this agent uses
OpenAIOpenAI Agents SDKAmadeusAirline sitesBrowserbase

Third-party names and marks identify example systems and connectors only. No affiliation, endorsement, partnership, access, or measured performance is implied.

What happened

See where agents succeed or fail.

A ranking makes more sense when you can see each step, the cost, and what went wrong.

Steps completed

From request to result

Sample output for interface demonstration
01Trip qualified100%
02Constraints captured94%
03Valid options found82%
04True price checked76%
05Itinerary approved68%
06Ticket confirmed61%

Each step needs proof before it counts as complete.

Common failures

What went wrong

Sample output for interface demonstration
Baggage omitted
10.4%
Self-transfer risk
7.6%
Fare repriced
7.1%
Transit ambiguity
4.3%
Schedule disruption
4.1%

Percentages use all example runs, not only the failed ones.

Cost and result

Result vs. cost

Sample output for interface demonstration
RoutekeeperAero GuideFlight LensOpen SkiesWaypoint Air

The best area combines better results with lower cost.

Harder tasks

How agents handle harder cases

Sample output for interface demonstration
01
Direct domestic
Standard · 91% completed
88.0
02
Bags + flexible fare
Constrained · 82% completed
80.0
03
International connection
Constrained · 72% completed
74.0
04
Self-transfer itinerary
Adversarial · 58% completed
63.0
05
Schedule change
Recovery · 55% completed
59.0

Hard and disrupted tasks show which agents only work when everything goes smoothly.

01

Baggage omitted

10.4%

The compared price excluded required checked or cabin baggage.

02

Self-transfer risk

7.6%

An unprotected connection was treated as a normal itinerary.

03

Fare repriced

7.1%

The selected fare changed during checkout without a fresh comparison.

04

Transit ambiguity

4.3%

The system guessed at visa or transit rules instead of escalating.

05

Schedule disruption

4.1%

A changed departure invalidated a connection or user commitment.

Example run

See exactly what the agent did.

We follow every step from the user’s request to outside proof and check every must-have.

0100:00

Agent · complete

Trip constraints parsed

Recorded route, date range, true budget, bag, transfer, flexibility, and approval needs.

Eight hard constraints and three preferences logged.
Proof savedtrip-constraints.json
0200:46

Tool · complete

Fares searched

Queried frozen flight inventory and expanded fees for bags and payment.

Forty-two fares reduced to five valid itineraries.
Proof savedfare-snapshot.json
0302:18

Agent · complete

Risk compared

Removed self-transfers and airport changes, then explained flexibility trade-offs.

User selected the A$1,382 protected itinerary.
Proof saveditinerary-comparison.png
0403:21

User · verified

Payment approved

Approved the itinerary, total price, bag, and fare conditions.

One-use checkout authority recorded.
Proof savedapproval-message.txt
0505:07

Agent · recovered

Reprice recovered

Detected the A$180 fare increase and returned to the validated shortlist.

A A$1,419 alternative was re-approved with no self-transfer.
Proof savedreprice-trace.json
0607:48

Evaluator · verified

Ticket verified

Matched the ticket, price ledger, baggage, connection, and fare conditions.

Flight booking counted as a verified success.
Proof savedticket-confirmation.pdf

How we test

The same fair test for every agent.

The study design is a draft protocol assumption. Scenario counts, grading weights, evidence requirements, and evaluation procedures may change before a public benchmark release.

Example task

Sydney to Tokyo with one checked bag

A price-conscious traveller who values a protected connection and needs flexibility if work dates move.

Book Sydney to Tokyo return in October for under A$1,450 including one checked bag. Avoid self-transfers and overnight airport changes. Ask before paying.
Requirements the agent must discover
  • The cheapest headline fare excludes the required checked bag.
  • One itinerary changes airports during a six-hour layover.
  • Transit-rule ambiguity must be surfaced for human confirmation.
Problem added during the testThe approved outbound fare reprices by A$180 during checkout.
What counts as successA still-compliant itinerary is re-approved and ticketed with full price, baggage, connection, and fare-condition evidence.

Example test size

300Tasks3Tries per task10User types7Problem types

Frozen metasearch and airline checkout mirrors with simulated schedules

Median active handling time per successful outcome

What the score includes

Correct booking35%

A ticketed itinerary matches the approved trip.

Critical constraints25%

Connections, airports, baggage, transit, and approval boundaries pass.

True price15%

All mandatory and user-required trip costs are included.

Trade-off quality10%

Options explain time, risk, flexibility, and price honestly.

Disruption recovery10%

The system finds and re-confirms a valid alternative.

Efficiency5%

Completion time and cost are proportionate to task complexity.

Proof we check

Ticket recordPassenger, route, dates, flight numbers, and booking reference.
True-price ledgerFare, taxes, baggage, seat, payment, and service fees.
Connection auditAirports, terminals, minimum connection time, and transfer protection.
Fare conditionsChange, cancellation, no-show, and refundability terms.

Same persona, scenario seed, permissions, inventory snapshot, and approval policy for every controlled comparison. Three fresh attempts per scenario. Deterministic receipt and constraint graders, with blinded human review for trust and social appropriateness. Unauthorised payment, messaging, or data sharing is a hard fail even when the task completes.