How we rate AI tools

A single benchmark never tells you which AI tool to buy. Our scores come from a weighted formula across the five factors that actually matter to a buyer — weighted, because capability matters more than documentation.

The formula

Every tool is rated 1–5 on each factor after hands-on testing. The overall score is the weighted sum:

Overall = (Capability × 35%) + (Value × 25%) + (Ease of use × 15%)
        + (Reliability × 15%) + (Support × 10%)
FactorWeightIn plain English
Capability35%How good the results actually are
Value for money25%What you get for the price
Ease of use15%How fast a beginner gets going
Speed & reliability15%Does it work every time, quickly
Docs & support10%Help when something goes wrong

Different jobs, different weights

The Finder goes one step further: when you pick a use case, we re-rank tools with weights tuned to that job. If you’re buying an AI to write code, raw capability is almost half the score; for an everyday assistant, ease of use leads.

Use caseCapabilityValue for moneyEase of useSpeed & reliabilityDocs & support
Write content40%20%20%10%10%
Write code45%15%10%20%10%
Make images45%20%20%10%5%
Make videos45%20%20%10%5%
Research & summarize40%15%15%25%5%
Study & learn30%25%30%10%5%
Everyday assistant25%25%35%10%5%
Audio & voice40%25%20%10%5%

How we test

  1. We use the tool for real work — not the marketing demo. Writing, coding, research, or image generation, depending on what the tool is for.
  2. We run the same set of prompts across competing tools in a category, so scores are comparable — and judge outputs on quality before checking which tool produced them.
  3. We watch the real-world numbers: response speed, how often it’s wrong or makes things up, whether the same prompt gives consistently good answers, and what it actually costs at realistic usage.
  4. We re-test after major updates — AI tools change monthly, and the “Updated” date on each review reflects the last time we checked.

Data sources

For model-level benchmarks (raw intelligence, token prices, latency across providers) we reference independent trackers such as Artificial Analysis alongside our own testing. Our scores are about the product experience a subscriber gets, which no raw benchmark fully captures.

Independence

Some links on review pages may earn us a commission. Commissions never influence scores, rankings, or verdicts — the weights above are fixed before we ever open a tool.