Main result
Recommendations acted on
- Acted on
- 43%
- Rank
- 1/12
- Range
- 40%–46%
- Runs
- 3,600
The user saved, shared, RSVP’d to, or booked the recommended event.
Real-world task test · EVT-05
Can an agent find something the user will actually want to do?
We test whether suggestions match the user, feel fresh, fit their time and budget, lead to action, and improve from feedback.
Highlights · EVT-05
Made-up example dataWe test the model, agent setup, and tools together. It only counts when we can verify the result.
Main result
The user saved, shared, RSVP’d to, or booked the recommended event.
Cost
The estimated model and tool cost for each completed task.
Speed
How long the agent actively worked on the task.
Complete agents
The user booked or RSVP’d, the event fit their calendar, and they attended.
Complete agents
A simple score combining task success, following requirements, recovery, cost, and trust.
Complete agents
How often the agent still finished after a price change, tool error, rejection, or timeout.
Complete agents
Whether the agent asked before important actions, showed its proof, and was honest about uncertainty.
Speed and cost
Best area is lower-left. Weekend Signal is closest to that area in this example, at 1m 36s fastest recommendation and $0.21 cheapest useful recommendation. Circle size shows recommendations acted on.
Agents closer to the lower-left finished faster and cost less. Bigger circles led to more actions.
Each model sees the same task and options. Tools are turned off, so the model cannot take actions.
Main result
How well the model understood the task and made a choice. The model could not use tools or take actions.
Models only
How often the model remembered the user’s must-haves when making a choice.
Models only
How good the model’s choice was when every model saw the same options.
Models only
How often the proposed steps were possible, in the right order, and safe to follow.
Models only
Whether the model used the available proof and avoided claiming success when it was unsure.
Speed and cost
Best area is upper-right. Gemini 2.5 Pro is closest to that area in this example, at 85% action plan works and 90% choice quality. Circle size shows honest about uncertainty.
Models in the upper-right make better choices and produce plans that are more likely to work. Bigger circles mean the model was more honest about uncertainty.
We keep the model and tools the same and change only the workflow that runs them.
Main result
How well the workflow ran the task when the model and tools stayed the same.
Agent setup
How often the setup followed the approved steps in the right order without losing progress.
Agent setup
How often the setup recovered from a timeout, outdated information, price change, or rejection.
Agent setup
Whether the setup remembered requirements, approvals, evidence, and completed steps.
Agent setup
How often the setup asked the user before an important or hard-to-reverse action.
Cost
The extra cost added by the agent setup for each completed task.
Speed and cost
Best area is upper-left. PydanticAI is closest to that area in this example, at $0.03 extra setup cost and 91% remembered progress. Circle size shows waited for approval.
Setups in the upper-left remember more while adding less cost. Bigger circles mean they waited for approval more reliably.
We send the same requests to each tool and check whether the action worked, the information was current, and proof came back.
Main result
How reliably the connected tool completed actions, returned current information, and provided proof.
Connected tools
How often the tool completed the requested action and returned proof.
Connected tools
How often the tool returned the correct price, availability, or policy at that moment.
Connected tools
How often the tool recovered after a temporary error.
Connected tools
How often the tool returned the details needed to prove what happened.
Speed and cost
Best area is upper-right. Eventbrite discovery connector is closest to that area in this example, at 95% information was current and 96% proof returned. Circle size shows retry worked.
Tools in the upper-right return more current information and better proof. Bigger circles mean retries worked more often.
Detailed comparison
Compare the complete agent, the model alone, the agent setup, and the connected tools.
Complete agent
We test the model, agent setup, and tools together. It only counts when we can verify the result.
Ranking
Higher is better. Small score differences may not matter once real test results replace these examples.
| Rank | Complete agents | Overall | Completed | Recovery | Trust | Cost | Time |
|---|---|---|---|---|---|---|---|
| 1 | Weekend SignalGemini 2.5 Pro · Google ADK | 75.9 | 42.0 | 72.0 | 87.0 | $0.21 | 1m 36s |
| 2 | Good PlansClaude Opus 4.1 · LangGraph | 74.7 | 40.0 | 75.0 | 91.0 | $0.27 | 1m 52s |
| 3 | Out ThereGPT-5 · OpenAI Agents SDK | 73.8 | 43.0 | 69.0 | 86.0 | $0.25 | 1m 43s |
| 4 | Local ThreadMistral Large 2 · PydanticAI | 68.6 | 36.0 | 65.0 | 81.0 | $0.17 | 2m 04s |
| 5 | Tonight MaybeGrok 4 · CrewAI | 64.9 | 34.0 | 59.0 | 75.0 | $0.23 | 2m 19s |
Third-party names and marks identify example systems and connectors only. No affiliation, endorsement, partnership, access, or measured performance is implied.
What happened
A ranking makes more sense when you can see each step, the cost, and what went wrong.
Steps completed
Each step needs proof before it counts as complete.
Common failures
Percentages use all example runs, not only the failed ones.
Cost and result
The best area combines better results with lower cost.
Harder tasks
Hard and disrupted tasks show which agents only work when everything goes smoothly.
The suggestion matched a broad category but not the person’s actual taste.
The recommendation was relevant but no longer actionable.
The event overlapped a known commitment or travel time.
Exploration drifted beyond the user’s tolerance without explanation.
The system resurfaced a declined event or ignored prior feedback.
Example run
We follow every step from the user’s request to outside proof and check every must-have.
Agent · complete
Combined the request with prior accepts, rejects, budget, access, and social comfort.
Three stable preferences and two exploration bounds recorded.Tool · complete
Queried local events and checked transit, calendar conflicts, price, and live capacity.
Fifty-six events reduced to six eligible candidates.Agent · complete
Presented one strong match and two alternatives with relevance and novelty reasons.
User opened the small gallery performance.Tool · complete
Rechecked availability before ticket handoff.
The preferred event had sold out.Agent · recovered
Promoted the pre-validated listening-room event and requested booking approval.
User approved a A$32 ticket and calendar entry.Evaluator · verified
Matched booking, calendar, simulated attendance, rating, and preference update.
Recommendation counted as a positive longitudinal outcome.How we test
The study design is a draft protocol assumption. Scenario counts, grading weights, evidence requirements, and evaluation procedures may change before a public benchmark release.
Example task
A 29-year-old new to Sydney who likes live music and design, avoids crowded clubs, has a A$45 budget, and wants to meet people without forced networking.
Find me something for Friday night that feels social but not awkward. I want to try something new, stay near a train line, and spend under A$45.
Example test size
Four simulated weeks of event catalogs, profiles, calendars, and feedback
Median active handling time per recommendation cycleWhat the score includes
The user saves, shares, books, or attends the recommendation.
The event matches stated and learned interests and constraints.
The system supports booking, calendar fit, reminders, and recovery.
Confidence and later recommendations improve with feedback.
The system expands taste without becoming irrelevant.
The system respects budget, access, privacy, and contact boundaries.
Proof we check
Same persona, scenario seed, permissions, inventory snapshot, and approval policy for every controlled comparison. Three fresh attempts per scenario. Deterministic receipt and constraint graders, with blinded human review for trust and social appropriateness. Unauthorised payment, messaging, or data sharing is a hard fail even when the task completes.