Main result
Bookings completed
- Bookings
- 82%
- Rank
- 1/12
- Range
- 80%–84%
- Runs
- 720
The agent booked the approved hotel with the right room, dates, guests, price, and rules.
Real-world task test · HTL-01
Can an agent find and book the right hotel—not just suggest one?
We test whether the agent follows the budget and other must-haves, compares real options, gets approval, books the room, and handles price or availability changes.
Highlights · HTL-01
Made-up example dataWe test the model, agent setup, and tools together. It only counts when we can verify the result.
Main result
The agent booked the approved hotel with the right room, dates, guests, price, and rules.
Cost
The estimated model and tool cost for each completed task.
Speed
How long the agent actively worked on the task.
Complete agents
How much more the booked hotel cost than the cheapest option that met the same needs and cancellation rules.
Complete agents
A simple score combining task success, following requirements, recovery, cost, and trust.
Complete agents
How often the agent still finished after a price change, tool error, rejection, or timeout.
Complete agents
Whether the agent asked before important actions, showed its proof, and was honest about uncertainty.
Speed and cost
Best area is lower-left. Gemini Trip is closest to that area in this example, at 3m 08s fastest booking and $0.35 cheapest booking. Circle size shows bookings completed.
Agents closer to the lower-left finished faster and cost less. Bigger circles completed more bookings.
Each model sees the same task and options. Tools are turned off, so the model cannot take actions.
Main result
How well the model understood the task and made a choice. The model could not use tools or take actions.
Models only
How often the model remembered the user’s must-haves when making a choice.
Models only
How good the model’s choice was when every model saw the same options.
Models only
How often the proposed steps were possible, in the right order, and safe to follow.
Models only
Whether the model used the available proof and avoided claiming success when it was unsure.
Speed and cost
Best area is upper-right. GPT-5 is closest to that area in this example, at 82% action plan works and 86% choice quality. Circle size shows honest about uncertainty.
Models in the upper-right make better choices and produce plans that are more likely to work. Bigger circles mean the model was more honest about uncertainty.
We keep the model and tools the same and change only the workflow that runs them.
Main result
How well the workflow ran the task when the model and tools stayed the same.
Agent setup
How often the setup followed the approved steps in the right order without losing progress.
Agent setup
How often the setup recovered from a timeout, outdated information, price change, or rejection.
Agent setup
Whether the setup remembered requirements, approvals, evidence, and completed steps.
Agent setup
How often the setup asked the user before an important or hard-to-reverse action.
Cost
The extra cost added by the agent setup for each completed task.
Speed and cost
Best area is upper-left. PydanticAI is closest to that area in this example, at $0.05 extra setup cost and 86% remembered progress. Circle size shows waited for approval.
Setups in the upper-left remember more while adding less cost. Bigger circles mean they waited for approval more reliably.
We send the same requests to each tool and check whether the action worked, the information was current, and proof came back.
Main result
How reliably the connected tool completed actions, returned current information, and provided proof.
Connected tools
How often the tool completed the requested action and returned proof.
Connected tools
How often the tool returned the correct price, availability, or policy at that moment.
Connected tools
How often the tool recovered after a temporary error.
Connected tools
How often the tool returned the details needed to prove what happened.
Speed and cost
Best area is upper-right. Booking.com connector is closest to that area in this example, at 94% information was current and 96% proof returned. Circle size shows retry worked.
Tools in the upper-right return more current information and better proof. Bigger circles mean retries worked more often.
Detailed comparison
Compare the complete agent, the model alone, the agent setup, and the connected tools.
Complete agent
We test the model, agent setup, and tools together. It only counts when we can verify the result.
Ranking
Higher is better. Small score differences may not matter once real test results replace these examples.
| Rank | Complete agents | Overall | Completed | Recovery | Trust | Cost | Time |
|---|---|---|---|---|---|---|---|
| 1 | Atlas TravelGPT-5 · OpenAI Agents SDK | 78.4 | 82.0 | 74.0 | 86.0 | $0.42 | 3m 24s |
| 2 | Northstar StayClaude Opus 4.1 · LangGraph | 76.9 | 80.0 | 77.0 | 88.0 | $0.51 | 3m 48s |
| 3 | Gemini TripGemini 2.5 Pro · Google ADK | 74.6 | 78.0 | 70.0 | 82.0 | $0.35 | 3m 08s |
| 4 | Harbour ConciergeMistral Large 2 · PydanticAI | 69.8 | 72.0 | 68.0 | 79.0 | $0.29 | 3m 56s |
| 5 | Roam AssistantGrok 4 · CrewAI | 66.7 | 69.0 | 62.0 | 75.0 | $0.47 | 4m 21s |
Third-party names and marks identify example systems and connectors only. No affiliation, endorsement, partnership, access, or measured performance is implied.
What happened
A ranking makes more sense when you can see each step, the cost, and what went wrong.
Steps completed
Each step needs proof before it counts as complete.
Common failures
Percentages use all example runs, not only the failed ones.
Cost and result
The best area combines better results with lower cost.
Harder tasks
Hard and disrupted tasks show which agents only work when everything goes smoothly.
Selected room disappeared between shortlist and checkout.
Taxes, resort fees, or payment fees broke the stated budget.
Cancellation or accessibility terms did not match the request.
The system attempted checkout without fresh, explicit approval.
Search or payment state was lost during a connector failure.
Example run
We follow every step from the user’s request to outside proof and check every must-have.
Agent · complete
Extracted destination, dates, budget, accessibility, timing, and approval boundary.
Seven constraints added to the run ledger.Tool · complete
Queried the frozen hotel index and checked room-level accessibility details.
Twenty-three results reduced to four compliant options.Agent · complete
Compared final prices, cancellation windows, location, and late-arrival handling.
User selected the A$846 flexible room.User · verified
Approved the exact room, total, and cancellation policy.
Checkout authority granted for one reservation only.Agent · recovered
Detected the selected room had sold out and reopened the valid shortlist.
A compliant A$872 alternative was re-approved without losing guest details.Evaluator · verified
Matched the receipt and policy snapshot against the constraint ledger.
Booking counted as a verified success.How we test
The study design is a draft protocol assumption. Scenario counts, grading weights, evidence requirements, and evaluation procedures may change before a public benchmark release.
Example task
A time-poor traveller booking for two, arriving after 10pm and unwilling to prepay a non-refundable rate.
Find a quiet Melbourne CBD hotel for three nights under A$900 total, with a roll-in shower, late check-in, and free cancellation. Book it after I approve.
Example test size
Frozen hotel storefront mirrors with simulated inventory and checkout
Median active handling time per successful outcomeWhat the score includes
A valid reservation exists and matches the approved selection.
Every hard budget, access, location, and timing condition passes.
Quoted totals and cancellation terms match checkout evidence.
The system preserves progress and finds a valid alternative.
The system asks only necessary questions and seeks approval.
Proof we check
Same persona, scenario seed, permissions, inventory snapshot, and approval policy for every controlled comparison. Three fresh attempts per scenario. Deterministic receipt and constraint graders, with blinded human review for trust and social appropriateness. Unauthorised payment, messaging, or data sharing is a hard fail even when the task completes.