Which AI model you should actually pay for
A Corolla, a Porsche and a Range Rover. A practical way to choose between AI tools, plus where the analogy breaks and how to test them on your own work.
Updated
Last reviewed for accuracy

A garage full of vehicles, and the question people keep asking wrong.
It would be convenient if one tool handled everything. It doesn’t, and the question of which one is best has no answer, because best is entirely a function of what you’re trying to do with it.
Here’s how I think about the three I use most. This is a snapshot rather than a law, since these things shift every few months, but the way of choosing outlasts the specifics.
ChatGPT is the Toyota Corolla. Rarely amazing, rarely terrible, gets you there. Reliable, unglamorous, does the job without drama. For standard daily work it’s the one that starts every time.
Claude is the Porsche. Sharpest through complex reasoning, and it has something closer to taste in how it writes. It also drinks fuel. You’ll burn through your allowance faster than you expect, so save it for the problems that genuinely need it.
Gemini is the Range Rover. Luxurious, very capable, handles rough terrain like video and images without complaining, though how much of that you get depends on which tier you’re on. Famously unreliable in the way Range Rovers are famously unreliable, which is to say it’ll be brilliant right up until the moment you needed it most.
Matching the vehicle to the road
Deep reasoning on something hard? Take the Porsche and watch the fuel gauge.
Standard daily work, the fifty small obligations that clog a week? The Corolla is right there.
Messy documents, video, images, anything where the input isn’t clean text? The Range Rover handles that terrain. Just don’t build a long trip around it.
Where the analogy stops being useful
Car metaphors are fun and they hide something, so it’s worth naming what they hide.
A car has fixed characteristics. These systems don’t. The rankings above will be wrong within a year, possibly within a quarter, because what changes is never the badge on the front. It’s what sits under the bonnet, and the company can replace it while you’re driving. Anyone confidently telling you a permanent hierarchy is describing a photograph and calling it a map.
That’s worth naming plainly, because the language everyone uses hides it. ChatGPT, Claude and Gemini are products. The models underneath them are separate things, with their own names and their own release schedule, and your subscription gives you whichever one is current. So when people argue about which model is best, they’re usually comparing products and calling them models. It sounds like a pedantic distinction and it isn’t, because it’s the difference between choosing something stable and choosing something that will be replaced without asking you.
The second thing the analogy hides is that most of the variance in output isn’t the model at all. It’s how well the problem was framed before it got handed over. I’ve watched people run identical tasks through three systems, get three mediocre results, and conclude the tools are overrated, when the actual issue was that the request contained no context about what good would look like. Switching cars doesn’t help if you haven’t decided where you’re going.
How to actually choose
Ignore the benchmarks and ignore me. Run your own test, which takes an afternoon and will be more useful than any comparison article.
Pick three tasks you genuinely do most weeks. Not impressive tasks, boring representative ones. The status summary, the first draft of the requirements, the awkward email. Run all three through each system with identical input. Then look at which output needed the least repair, because repair time is the only cost that matters and it’s the one nobody measures.
You’ll usually find the differences are smaller than the discourse suggests, and that they cluster in a specific place. One will be better at your particular kind of work, for reasons neither of you could fully articulate, and that’s enough to decide on.
The point underneath all of it
People ask which one is best because they’re hoping to make the decision once and stop thinking about it. That’s a reasonable thing to want and it’s the wrong shape for this problem.
Most people I know are paying for two or three subscriptions and using one of them for ninety percent of everything. That’s not a tooling problem. It’s an unexamined-defaults problem, and it costs about the same as a gym membership you don’t use.
Tagged
- ChatGPT
- Claude
- Gemini
- AI tools
- comparison
- AI and Work
