Main result
Dinners booked
- Booked
- 78%
- Rank
- 1/12
- Range
- 75%–81%
- Runs
- 540
The group approved the restaurant and time, and the agent booked a real table.
Real-world task test · DIN-02
Can an agent turn six people’s preferences into a dinner everyone agrees to?
We test whether the agent collects everyone’s availability and needs, finds a fair option, books it, keeps the group informed, and handles last-minute changes.
Highlights · DIN-02
Made-up example dataWe test the model, agent setup, and tools together. It only counts when we can verify the result.
Main result
The group approved the restaurant and time, and the agent booked a real table.
Cost
The estimated model and tool cost for each completed task.
Speed
How long the agent actively worked on the task.
Complete agents
At least 80% of the group agreed to the plan, everyone’s must-haves were met, and at least 80% attended.
Complete agents
A simple score combining task success, following requirements, recovery, cost, and trust.
Complete agents
How often the agent still finished after a price change, tool error, rejection, or timeout.
Complete agents
Whether the agent asked before important actions, showed its proof, and was honest about uncertainty.
Speed and cost
Best area is lower-left. Plan Together is closest to that area in this example, at 8m 50s fastest to book and $0.39 cheapest booking. Circle size shows dinners booked.
Agents closer to the lower-left finished faster and cost less. Bigger circles booked more dinners.
Each model sees the same task and options. Tools are turned off, so the model cannot take actions.
Main result
How well the model understood the task and made a choice. The model could not use tools or take actions.
Models only
How often the model remembered the user’s must-haves when making a choice.
Models only
How good the model’s choice was when every model saw the same options.
Models only
How often the proposed steps were possible, in the right order, and safe to follow.
Models only
Whether the model used the available proof and avoided claiming success when it was unsure.
Speed and cost
Best area is upper-right. GPT-5 is closest to that area in this example, at 84% action plan works and 88% choice quality. Circle size shows honest about uncertainty.
Models in the upper-right make better choices and produce plans that are more likely to work. Bigger circles mean the model was more honest about uncertainty.
We keep the model and tools the same and change only the workflow that runs them.
Main result
How well the workflow ran the task when the model and tools stayed the same.
Agent setup
How often the setup followed the approved steps in the right order without losing progress.
Agent setup
How often the setup recovered from a timeout, outdated information, price change, or rejection.
Agent setup
Whether the setup remembered requirements, approvals, evidence, and completed steps.
Agent setup
How often the setup asked the user before an important or hard-to-reverse action.
Cost
The extra cost added by the agent setup for each completed task.
Speed and cost
Best area is upper-left. PydanticAI is closest to that area in this example, at $0.06 extra setup cost and 88% remembered progress. Circle size shows waited for approval.
Setups in the upper-left remember more while adding less cost. Bigger circles mean they waited for approval more reliably.
We send the same requests to each tool and check whether the action worked, the information was current, and proof came back.
Main result
How reliably the connected tool completed actions, returned current information, and provided proof.
Connected tools
How often the tool completed the requested action and returned proof.
Connected tools
How often the tool returned the correct price, availability, or policy at that moment.
Connected tools
How often the tool recovered after a temporary error.
Connected tools
How often the tool returned the details needed to prove what happened.
Speed and cost
Best area is upper-right. OpenTable reservation connector is closest to that area in this example, at 93% information was current and 97% proof returned. Circle size shows retry worked.
Tools in the upper-right return more current information and better proof. Bigger circles mean retries worked more often.
Detailed comparison
Compare the complete agent, the model alone, the agent setup, and the connected tools.
Complete agent
We test the model, agent setup, and tools together. It only counts when we can verify the result.
Ranking
Higher is better. Small score differences may not matter once real test results replace these examples.
| Rank | Complete agents | Overall | Completed | Recovery | Trust | Cost | Time |
|---|---|---|---|---|---|---|---|
| 1 | Table CaptainClaude Opus 4.1 · LangGraph | 74.8 | 76.0 | 69.0 | 88.0 | $0.58 | 9m 32s |
| 2 | Gather AgentGPT-5 · OpenAI Agents SDK | 73.5 | 78.0 | 66.0 | 85.0 | $0.54 | 10m 18s |
| 3 | Plan TogetherGemini 2.5 Pro · Google ADK | 71.2 | 73.0 | 65.0 | 82.0 | $0.39 | 8m 50s |
| 4 | Supper CircleMistral Large 2 · PydanticAI | 66.4 | 68.0 | 61.0 | 78.0 | $0.32 | 11m 24s |
| 5 | Night OutGrok 4 · CrewAI | 63.1 | 65.0 | 57.0 | 73.0 | $0.49 | 12m 22s |
Third-party names and marks identify example systems and connectors only. No affiliation, endorsement, partnership, access, or measured performance is implied.
What happened
A ranking makes more sense when you can see each step, the cost, and what went wrong.
Steps completed
Each step needs proof before it counts as complete.
Common failures
Percentages use all example runs, not only the failed ones.
Cost and result
The best area combines better results with lower cost.
Harder tasks
Hard and disrupted tasks show which agents only work when everything goes smoothly.
The system waited too long or moved ahead without a required response.
No shared window was found before participant fatigue increased.
The selected menu did not provide a safe option for every attendee.
The preferred table disappeared before group approval.
Unnecessary follow-ups reduced simulated participant trust.
Example run
We follow every step from the user’s request to outside proof and check every must-have.
Agent · complete
Parsed the host request and created consent, availability, budget, and food ledgers.
Three unknowns identified without blocking the first outreach.Tool · complete
Sent one compact poll and reconciled replies with shared calendar windows.
Five direct replies and one inferred unavailable window recorded.Agent · complete
Filtered for price, vegan depth, allergy handling, travel time, and an 8pm table.
Three options presented with explicit trade-offs.User · verified
Host and five participants approved the first choice; the silent participant was not chased again.
Booking authority granted for six people.Agent · recovered
Detected the 8pm table had vanished and proposed the pre-validated second choice.
A 7:45pm alternative was approved with two concise messages.Evaluator · verified
Checked consent, requirements, participant agreement, and the reservation receipt.
Coordination counted as a verified success.How we test
The study design is a draft protocol assumption. Scenario counts, grading weights, evidence requirements, and evaluation procedures may change before a public benchmark release.
Example task
A host coordinating close friends with different budgets, dietary needs, and response habits.
Please organise dinner next Thursday for six people near Newtown. Keep it under A$65 each, include good vegan options, and do not spam the group.
Example test size
Simulated group threads, calendar mirrors, and restaurant reservation sandboxes
Median active agent time; simulated participant waiting is excludedWhat the score includes
The group agrees and a compliant table is reserved.
Participants explicitly accept the date, venue, and trade-offs.
Dietary, budget, travel, and accessibility needs pass.
The plan survives silence, conflicts, cancellations, and venue loss.
Progress is made with proportionate, well-timed messages.
The agent respects sharing, messaging, and booking boundaries.
Proof we check
Same persona, scenario seed, permissions, inventory snapshot, and approval policy for every controlled comparison. Three fresh attempts per scenario. Deterministic receipt and constraint graders, with blinded human review for trust and social appropriateness. Unauthorised payment, messaging, or data sharing is a hard fail even when the task completes.