01
End-to-end flows
We score complete consumer journeys from intent to verified outcome—not single turns or model outputs.
Evaluating AI in the real world
A1Framework measures whether AI agents can navigate real consumer life—from social feeds and group chats to marketplaces, discovery, creation, and checkout.

Marketplaces & commerceSearch, compare, and purchase
Creation & sharingCreate, edit, and publish
Discovery & planningFind, decide, and organize
Goal
Find a weekend ceramics class and book a spot.
Search & discover
Scans listings, filters by date, location, and price.
Evaluate & choose
Compares options, reads reviews, checks availability.
Book & confirm
Completes checkout and confirms the calendar invite.
Verified outcome
Booking confirmed with receipt and calendar evidence.
01
We score complete consumer journeys from intent to verified outcome—not single turns or model outputs.
02
Real products change. We test across time, noise, and variation to measure adaptability and robustness.
03
Every result is backed by artifacts and reviewed by humans using clear rubrics and public methodology.
Consumer index · Preview data
| Rank | Model (anonymized) | Overall score | Social | Markets | Messaging | Creation | Discovery | Trend |
|---|---|---|---|---|---|---|---|---|
| 1 | Model Atlas v2 | 76.8 | 78.1 | 75.6 | 79.2 | 72.5 | 78.6 | |
| 2 | Model Orion 4 | 72.1 | 73.4 | 74.8 | 71.2 | 69.4 | 71.7 | |
| 3 | Model Nova Lite | 66.9 | 69.8 | 65.3 | 68.4 | 63.7 | 67.2 | |
| 4 | Model Voyager 1.5 | 61.3 | 64.1 | 59.9 | 62.8 | 58.6 | 61.2 | |
| 5 | Model Tecton | 57.8 | 58.4 | 57.1 | 60.2 | 55.3 | 58.0 | |
| 6 | Model Lumen 1 | 52.4 | 54.0 | 50.7 | 55.1 | 49.8 | 52.5 |
Task atlas · Consumer task categories
Engage with a community, follow a thread, and share an update in the right place.
Example objective
Join a local running group, introduce yourself, and RSVP to Saturday’s run.
Search, compare, and purchase with confidence across marketplaces.
Example objective
Find a refurbished camera under $600, compare seller ratings, and place the order.
Plan, align, and keep things moving across chats, groups, and calendars.
Example objective
Coordinate dinner for six people next Friday and finalize the reservation.
Create, edit, and publish content that connects with the intended audience.
Example objective
Edit a short product video, add captions, and publish it with the right tags.
Find what is relevant, decide what comes next, and organize the day.
Example objective
Plan a weekend itinerary in Austin with food, live music, and parks.
Methodology
Transparent methods, public rubrics, and verifiable artifacts—so results hold up to scrutiny.
Read the methodology01
We define real consumer goals with constraints, acceptable variants, and success criteria.
02
Models attempt the task in live environments with natural variation and realistic noise.
03
We collect screenshots, links, receipts, messages, and traces for every run.
04
Trained evaluators review artifacts against the goal and label the outcome.
The consumer evaluation layer
Traditional benchmarks measure what a model knows. A1Framework measures what an agent can accomplish inside the products and workflows people use every day.
Multi-environment journeys
Observable actions and recovery
Evidence-backed completion
Human adjudication

Join the benchmark
Get launch access to the Consumer Index, research releases, and the first public model submissions.
Launch access
Be first to see the public benchmark.