Private questions. Human and AI review.
How we evaluate.
An answer should do more than sound right.
Here’s how we evaluate its usefulness in everyday life.
01 / Inside the evaluation
What makes an answer useful?
Choose a category. Follow the answer
all the way through its review.
It’s Tuesday. I’m moving Saturday, but only have an hour after work each evening. How should I start?
Response AIdentity hidden
Start by making a moving checklist and gathering boxes. Pack one room at a time, labeling everything clearly.
Remember to update your address and confirm your movers. Keep essential items handy for your first night.
Response BIdentity hidden
Tonight: spend 15 minutes confirming transport, then 45 packing books and decor.
Wednesday: pack off-season clothes. Thursday: kitchen extras, leaving one set of dishes.
Friday: pack bedding and a first-night bag. Keep keys, medication, and chargers with you.
02 / From reviews to rankings
Two perspectives. Equal weight.
Results accumulate across the question set.
Each category contributes equally to the overall score.
Illustrative aggregate scores · not a real model’s results
A category score
Human review + AI judging
84 / 100
92 / 100
The two ratings are shown on the same 100-point scale. Each supplies half of this category’s score.
The overall score
Five categories, one average
- General Helpfulness
- 88
- Emotional Support
- 86
- Creative Writing
- 90
- Reading Comprehension
- 94
- Instruction Following
- 92
(88 + 86 + 90 + 94 + 92) ÷ 5
Speed and price are separate. They don’t change this score.
03 / The question set
Private questions.
Open methodology.
We keep the test questions private to reduce the risk of models memorizing the benchmark. We share how we evaluate them so you can understand what the scores mean.
The questions stay private
Prompts are created for everyday situations and reviewed by people. The questions on this page were written separately for the walkthrough.
The process is explained
Five dimensions of assistant usefulness. Human review and AI judging, weighted equally. Task-specific checks for instructions and comprehension.
Dataset scope & version details
The v1.0 system card documents approximately 500 synthetic, human-reviewed prompts created in July 2025. It focuses on English, single-turn and short-context tasks.
It names Claude 4.5 Sonnet and GPT-5.1 as writing judges and describes human review alongside programmatic checks. Those are version-specific details; this page explains the current 50/50 human and AI weighting.
Read the v1.0 system cardA few more details
Questions about the process
Are these the actual benchmark questions?
No. The walkthrough uses newly written questions, prepared answers, and illustrative reviews. The benchmark question set stays private to reduce contamination and overfitting.
Who reviews the answers?
Human reviewers and AI judges both contribute to evaluation, with a 50/50 weighting. Human review considers usefulness and nuance. AI judging adds a structured assessment. Task-specific checks examine constraints and facts in supplied passages.
How is the overall score calculated?
Human review and AI judging contribute equally to category scoring. The overall score is the average of the five category scores: General Helpfulness, Emotional Support, Creative Writing, Reading Comprehension, and Instruction Following. Speed and price are shown separately.
What does the benchmark cover?
Chatio focuses on everyday assistance in English, using single-turn and short-context tasks. It measures practical usefulness, tone, writing, comprehension, and instruction following. It does not evaluate the full range of coding, advanced mathematics, or multimodal capabilities.
Find your everyday assistant