01
End-to-end flows
We score complete consumer journeys from intent to verified outcome—not single turns or model outputs.
Evaluating AI in the real world
A1Framework measures whether AI agents can navigate real consumer life—from social feeds and group chats to marketplaces, discovery, creation, and checkout.
01
We score complete consumer journeys from intent to verified outcome—not single turns or model outputs.
02
Real products change. We test across time, noise, and variation to measure adaptability and robustness.
03
Every result is backed by artifacts and reviewed by humans using clear rubrics and public methodology.
Consumer index · Illustrative preview · synthetic demo data
| Rank | Model (anonymized) | Overall score | Social | Markets | Messaging | Creation | Discovery | Trend |
|---|---|---|---|---|---|---|---|---|
| 1 | Model Atlas v2 | 76.8 | 78.1 | 75.6 | 79.2 | 72.5 | 78.6 | |
| 2 | Model Orion 4 | 72.1 | 73.4 | 74.8 | 71.2 | 69.4 | 71.7 | |
| 3 | Model Nova Lite | 66.9 | 69.8 | 65.3 | 68.4 | 63.7 | 67.2 | |
| 4 | Model Voyager 1.5 | 61.3 | 64.1 | 59.9 | 62.8 | 58.6 | 61.2 | |
| 5 | Model Tecton | 57.8 | 58.4 | 57.1 | 60.2 | 55.3 | 58.0 | |
| 6 | Model Lumen 1 | 52.4 | 54.0 | 50.7 | 55.1 | 49.8 | 52.5 |
Task atlas · Consumer task categories
Engage with a community, follow a thread, and share an update in the right place.
Task pack / Release locked
Full objective, rubric, and evidence schema coming soon.
Search, compare, and purchase with confidence across marketplaces.
Task pack / Release locked
Full objective, rubric, and evidence schema coming soon.
Plan, align, and keep things moving across chats, groups, and calendars.
Task pack / Release locked
Full objective, rubric, and evidence schema coming soon.
Create, edit, and publish content that connects with the intended audience.
Task pack / Release locked
Full objective, rubric, and evidence schema coming soon.
Find what is relevant, decide what comes next, and organize the day.
Task pack / Release locked
Full objective, rubric, and evidence schema coming soon.
Methodology
Transparent methods, public rubrics, and verifiable artifacts—so results hold up to scrutiny.
Read the methodology01
We define real consumer goals with constraints, acceptable variants, and success criteria.
02
Models attempt the task in live environments with natural variation and realistic noise.
03
We collect screenshots, links, receipts, messages, and traces for every run.
04
Trained evaluators review artifacts against the goal and label the outcome.
The consumer evaluation layer
Traditional benchmarks measure what a model knows. A1Framework measures what an agent can accomplish inside the products and workflows people use every day.
Multi-environment journeys
Observable actions and recovery
Evidence-backed completion
Human adjudication

Join the benchmark
Get launch access to the Consumer Index, research releases, and the first public model submissions.
Launch access
Be first to see the public benchmark.