Evaluating AI in the real world

Consumer AI Benchmarks

A1Framework measures whether AI agents can navigate real consumer life—from social feeds and group chats to marketplaces, discovery, creation, and checkout.

Human-verifiedOutcome-basedBuilt for consumer AI
A connected map of consumer task categories

Social & communityEngage, share, and belong

Marketplaces & commerceSearch, compare, and purchase

Messaging & coordinationPlan, align, and get things done

Creation & sharingCreate, edit, and publish

Discovery & planningFind, decide, and organize

Single run traceIllustrative preview

Goal

Find a weekend ceramics class and book a spot.

Search & discover

Scans listings, filters by date, location, and price.

Evaluate & choose

Compares options, reads reviews, checks availability.

Book & confirm

Completes checkout and confirms the calendar invite.

Verified outcome

Booking confirmed with receipt and calendar evidence.

We evaluate outcomes, not isolated answers.

01

End-to-end flows

We score complete consumer journeys from intent to verified outcome—not single turns or model outputs.

02

Dynamic environments

Real products change. We test across time, noise, and variation to measure adaptability and robustness.

03

Human-verified evidence

Every result is backed by artifacts and reviewed by humans using clear rubrics and public methodology.

Consumer index · Preview data

How models perform on real consumer tasks.

Preview releaseJuly 28, 2026
Rank by
RankModel (anonymized)Overall scoreSocialMarketsMessagingCreationDiscoveryTrend
1Model Atlas v276.878.175.679.272.578.6
2Model Orion 472.173.474.871.269.471.7
3Model Nova Lite66.969.865.368.463.767.2
4Model Voyager 1.561.364.159.962.858.661.2
5Model Tecton57.858.457.160.255.358.0
6Model Lumen 152.454.050.755.149.852.5
1
Model

Model Atlas v2

Overall score76.8
Social78.1
Markets75.6
Messaging79.2
Creation72.5
Discovery78.6
2
Model

Model Orion 4

Overall score72.1
Social73.4
Markets74.8
Messaging71.2
Creation69.4
Discovery71.7
3
Model

Model Nova Lite

Overall score66.9
Social69.8
Markets65.3
Messaging68.4
Creation63.7
Discovery67.2
4
Model

Model Voyager 1.5

Overall score61.3
Social64.1
Markets59.9
Messaging62.8
Creation58.6
Discovery61.2
5
Model

Model Tecton

Overall score57.8
Social58.4
Markets57.1
Messaging60.2
Creation55.3
Discovery58.0
6
Model

Model Lumen 1

Overall score52.4
Social54.0
Markets50.7
Messaging55.1
Creation49.8
Discovery52.5

Task atlas · Consumer task categories

Benchmarked where life happens.

Explore how tasks are built

Social & community

Engage with a community, follow a thread, and share an update in the right place.

Example objective

Join a local running group, introduce yourself, and RSVP to Saturday’s run.

Complexity

Marketplaces & commerce

Search, compare, and purchase with confidence across marketplaces.

Example objective

Find a refurbished camera under $600, compare seller ratings, and place the order.

Complexity

Messaging & coordination

Plan, align, and keep things moving across chats, groups, and calendars.

Example objective

Coordinate dinner for six people next Friday and finalize the reservation.

Complexity

Creation & sharing

Create, edit, and publish content that connects with the intended audience.

Example objective

Edit a short product video, add captions, and publish it with the right tags.

Complexity

Discovery & planning

Find what is relevant, decide what comes next, and organize the day.

Example objective

Plan a weekend itinerary in Austin with food, live music, and parks.

Complexity

Methodology

Every score has a receipt.

Transparent methods, public rubrics, and verifiable artifacts—so results hold up to scrutiny.

Read the methodology

01

Task specification

We define real consumer goals with constraints, acceptable variants, and success criteria.

02

Agent execution

Models attempt the task in live environments with natural variation and realistic noise.

03

Artifact capture

We collect screenshots, links, receipts, messages, and traces for every run.

04

Human verification

Trained evaluators review artifacts against the goal and label the outcome.

The consumer evaluation layer

The missing measurement layer for consumer agents.

Traditional benchmarks measure what a model knows. A1Framework measures what an agent can accomplish inside the products and workflows people use every day.

Multi-environment journeys

Observable actions and recovery

Evidence-backed completion

Human adjudication

Join the benchmark

Evaluate the next generation of consumer AI.

Get launch access to the Consumer Index, research releases, and the first public model submissions.

Launch access

Be first to see the public benchmark.

Interested in submitting a model?