Skip to methodology

Private questions. Human and AI review.

How we evaluate.

An answer should do more than sound right.
Here’s how we evaluate its usefulness in everyday life.

01 / Inside the evaluation

What makes an answer useful?

Choose a category. Follow the answer
all the way through its review.

Illustrative example · not a benchmark questionPrepared answers & reviews
chatioBlind comparison
Question

It’s Tuesday. I’m moving Saturday, but only have an hour after work each evening. How should I start?

Response AIdentity hidden

Start by making a moving checklist and gathering boxes. Pack one room at a time, labeling everything clearly.

Remember to update your address and confirm your movers. Keep essential items handy for your first night.

Response BIdentity hidden

Tonight: spend 15 minutes confirming transport, then 45 packing books and decor.

Wednesday: pack off-season clothes. Thursday: kitchen extras, leaving one set of dishes.

Friday: pack bedding and a first-night bag. Keep keys, medication, and chargers with you.

Example completeA short look at how evaluation works
One example explains the process. Many questions make a benchmark.How scores combine

02 / From reviews to rankings

Two perspectives. Equal weight.

Results accumulate across the question set.
Each category contributes equally to the overall score.

Illustrative aggregate scores · not a real model’s results

A category score

Human review + AI judging

Human review

84 / 100

50% weight 42 points
AI judging

92 / 100

50% weight 46 points
424688 / 100

The two ratings are shown on the same 100-point scale. Each supplies half of this category’s score.

The overall score

Five categories, one average

General Helpfulness
88
Emotional Support
86
Creative Writing
90
Reading Comprehension
94
Instruction Following
92
Overall score

(88 + 86 + 90 + 94 + 92) ÷ 5

90 / 100

Speed and price are separate. They don’t change this score.

Compare the overall score, then explore the dimensions that matter to you.

03 / The question set

Private questions.
Open methodology.

We keep the test questions private to reduce the risk of models memorizing the benchmark. We share how we evaluate them so you can understand what the scores mean.

The questions stay private

Prompts are created for everyday situations and reviewed by people. The questions on this page were written separately for the walkthrough.

The process is explained

Five dimensions of assistant usefulness. Human review and AI judging, weighted equally. Task-specific checks for instructions and comprehension.

Dataset scope & version details

The v1.0 system card documents approximately 500 synthetic, human-reviewed prompts created in July 2025. It focuses on English, single-turn and short-context tasks.

It names Claude 4.5 Sonnet and GPT-5.1 as writing judges and describes human review alongside programmatic checks. Those are version-specific details; this page explains the current 50/50 human and AI weighting.

Read the v1.0 system card

A few more details

Questions about the process

Are these the actual benchmark questions?

No. The walkthrough uses newly written questions, prepared answers, and illustrative reviews. The benchmark question set stays private to reduce contamination and overfitting.

Who reviews the answers?

Human reviewers and AI judges both contribute to evaluation, with a 50/50 weighting. Human review considers usefulness and nuance. AI judging adds a structured assessment. Task-specific checks examine constraints and facts in supplied passages.

How is the overall score calculated?

Human review and AI judging contribute equally to category scoring. The overall score is the average of the five category scores: General Helpfulness, Emotional Support, Creative Writing, Reading Comprehension, and Instruction Following. Speed and price are shown separately.

What does the benchmark cover?

Chatio focuses on everyday assistance in English, using single-turn and short-context tasks. It measures practical usefulness, tone, writing, comprehension, and instruction following. It does not evaluate the full range of coding, advanced mathematics, or multimodal capabilities.

Find your everyday assistant

See how the models compare.

Explore the leaderboard