How we rate AI tools
A single benchmark never tells you which AI tool to buy. Our scores come from a weighted formula across the five factors that actually matter to a buyer — weighted, because capability matters more than documentation.
The formula
Every tool is rated 1–5 on each factor after hands-on testing. The overall score is the weighted sum:
Overall = (Capability × 35%) + (Value × 25%) + (Ease of use × 15%)
+ (Reliability × 15%) + (Support × 10%)| Factor | Weight | In plain English |
|---|---|---|
| Capability | 35% | How good the results actually are |
| Value for money | 25% | What you get for the price |
| Ease of use | 15% | How fast a beginner gets going |
| Speed & reliability | 15% | Does it work every time, quickly |
| Docs & support | 10% | Help when something goes wrong |
Different jobs, different weights
The Finder goes one step further: when you pick a use case, we re-rank tools with weights tuned to that job. If you’re buying an AI to write code, raw capability is almost half the score; for an everyday assistant, ease of use leads.
| Use case | Capability | Value for money | Ease of use | Speed & reliability | Docs & support |
|---|---|---|---|---|---|
| Write content | 40% | 20% | 20% | 10% | 10% |
| Write code | 45% | 15% | 10% | 20% | 10% |
| Make images | 45% | 20% | 20% | 10% | 5% |
| Make videos | 45% | 20% | 20% | 10% | 5% |
| Research & summarize | 40% | 15% | 15% | 25% | 5% |
| Study & learn | 30% | 25% | 30% | 10% | 5% |
| Everyday assistant | 25% | 25% | 35% | 10% | 5% |
| Audio & voice | 40% | 25% | 20% | 10% | 5% |
How we test
- We use the tool for real work — not the marketing demo. Writing, coding, research, or image generation, depending on what the tool is for.
- We run the same set of prompts across competing tools in a category, so scores are comparable — and judge outputs on quality before checking which tool produced them.
- We watch the real-world numbers: response speed, how often it’s wrong or makes things up, whether the same prompt gives consistently good answers, and what it actually costs at realistic usage.
- We re-test after major updates — AI tools change monthly, and the “Updated” date on each review reflects the last time we checked.
Data sources
For model-level benchmarks (raw intelligence, token prices, latency across providers) we reference independent trackers such as Artificial Analysis alongside our own testing. Our scores are about the product experience a subscriber gets, which no raw benchmark fully captures.
Independence
Some links on review pages may earn us a commission. Commissions never influence scores, rankings, or verdicts — the weights above are fixed before we ever open a tool.