Evaluating AI in the real world

Consumer AI Benchmarks

A1Framework measures whether AI agents can navigate real consumer life—from social feeds and group chats to marketplaces, discovery, creation, and checkout.

Human-verifiedOutcome-basedBuilt for consumer AI
A1F / Consumer task field05 domains · live topology
Node 01 · SOCSocial & communityEngage, share, and belong
Node 02 · COMMarketplaces & commerceSearch, compare, and purchase
Node 03 · MSGMessaging & coordinationPlan, align, and get things done
Node 04 · CRTCreation & sharingCreate, edit, and publish
Node 05 · DSCDiscovery & planningFind, decide, and organize

Field data / Release locked

Live evaluation traces are compiling.

Verified actions, artifacts, and outcome receipts will publish with the first Consumer Index drop.

Calibration active

We evaluate outcomes, not isolated answers.

01

End-to-end flows

We score complete consumer journeys from intent to verified outcome—not single turns or model outputs.

02

Dynamic environments

Real products change. We test across time, noise, and variation to measure adaptability and robustness.

03

Human-verified evidence

Every result is backed by artifacts and reviewed by humans using clear rubrics and public methodology.

Consumer index · Illustrative preview · synthetic demo data

How models perform on real consumer tasks.

Preview statusIllustrative only
Rank by
RankModel (anonymized)Overall scoreSocialMarketsMessagingCreationDiscoveryTrend
1Model Atlas v276.878.175.679.272.578.6
2Model Orion 472.173.474.871.269.471.7
3Model Nova Lite66.969.865.368.463.767.2
4Model Voyager 1.561.364.159.962.858.661.2
5Model Tecton57.858.457.160.255.358.0
6Model Lumen 152.454.050.755.149.852.5
1
Model

Model Atlas v2

Overall score76.8
Social78.1
Markets75.6
Messaging79.2
Creation72.5
Discovery78.6
2
Model

Model Orion 4

Overall score72.1
Social73.4
Markets74.8
Messaging71.2
Creation69.4
Discovery71.7
3
Model

Model Nova Lite

Overall score66.9
Social69.8
Markets65.3
Messaging68.4
Creation63.7
Discovery67.2
4
Model

Model Voyager 1.5

Overall score61.3
Social64.1
Markets59.9
Messaging62.8
Creation58.6
Discovery61.2
5
Model

Model Tecton

Overall score57.8
Social58.4
Markets57.1
Messaging60.2
Creation55.3
Discovery58.0
6
Model

Model Lumen 1

Overall score52.4
Social54.0
Markets50.7
Messaging55.1
Creation49.8
Discovery52.5

Task atlas · Consumer task categories

Benchmarked where life happens.

Explore how tasks are built

Social & community

Engage with a community, follow a thread, and share an update in the right place.

Task pack / Release locked

Full objective, rubric, and evidence schema coming soon.

Atlas compiling

Marketplaces & commerce

Search, compare, and purchase with confidence across marketplaces.

Task pack / Release locked

Full objective, rubric, and evidence schema coming soon.

Atlas compiling

Messaging & coordination

Plan, align, and keep things moving across chats, groups, and calendars.

Task pack / Release locked

Full objective, rubric, and evidence schema coming soon.

Atlas compiling

Creation & sharing

Create, edit, and publish content that connects with the intended audience.

Task pack / Release locked

Full objective, rubric, and evidence schema coming soon.

Atlas compiling

Discovery & planning

Find what is relevant, decide what comes next, and organize the day.

Task pack / Release locked

Full objective, rubric, and evidence schema coming soon.

Atlas compiling

Methodology

Every score has a receipt.

Transparent methods, public rubrics, and verifiable artifacts—so results hold up to scrutiny.

Read the methodology

01

Task specification

We define real consumer goals with constraints, acceptable variants, and success criteria.

02

Agent execution

Models attempt the task in live environments with natural variation and realistic noise.

03

Artifact capture

We collect screenshots, links, receipts, messages, and traces for every run.

04

Human verification

Trained evaluators review artifacts against the goal and label the outcome.

The consumer evaluation layer

The missing measurement layer for consumer agents.

Traditional benchmarks measure what a model knows. A1Framework measures what an agent can accomplish inside the products and workflows people use every day.

Multi-environment journeys

Observable actions and recovery

Evidence-backed completion

Human adjudication

Join the benchmark

Evaluate the next generation of consumer AI.

Get launch access to the Consumer Index, research releases, and the first public model submissions.

Launch access

Be first to see the public benchmark.

Interested in submitting a model?