Guide · AI Assistants
How to choose an AI assistant for real work
A decision process for picking a general-purpose AI assistant based on the work you actually do, not on feature lists or benchmark scores.

General-purpose AI assistants have converged. The major ones all hold a conversation, read documents you give them, write in a requested style, and reason through a multi-step question. Feature comparison tables have become close to useless for choosing between them, because the differences that matter for a given person are rarely the differences a table shows.
The useful question is not which assistant is best. It is which assistant is best at the specific thing you will ask it to do fifty times a month. That question is answerable in an afternoon, and the process below is how to answer it without relying on anyone's marketing.
Step one: name the job in one sentence
Before opening a single product page, write the task down concretely. "Help with writing" is not a task. "Turn my rough meeting notes into a structured summary a client can read" is a task. "Read a forty-page supplier contract and tell me what obligations it puts on us" is a task. The sentence should name the input, the output, and who reads the output.
This sentence does most of the work of choosing. An assistant that excels at long-document reasoning is not necessarily the one that writes the most natural short marketing copy. An assistant integrated into the office suite you already live in may beat a technically stronger one that lives in a separate browser tab you forget to open.
Step two: assemble a small, honest test set
Take three real examples of the input from your own work. Choose one easy case, one typical case, and one genuinely awkward case — the messy notes, the badly scanned document, the request that contradicts itself. The awkward case is the important one. Every assistant looks capable on clean input; they separate on difficult input.
Run the same three cases through each candidate with the same instruction. Save the outputs somewhere you can compare side by side. This is the whole evaluation, and it is far more informative than any published benchmark, because benchmarks measure average performance on other people's problems.
Step three: score what you will actually notice
- Accuracy on your material: does it invent details, or does it stay inside what you gave it?
- Instruction-following: when you ask for four bullet points in plain English, do you get four bullet points in plain English?
- Recovery: when the first answer is wrong, how quickly does a correction get you to a usable one?
- Effort to usable: how much editing stands between the output and something you would send?
- Consistency: run the awkward case twice. Wildly different answers are a reliability problem.
Notice that speed is missing from that list. For most knowledge work, the difference between a four-second and a twelve-second response is irrelevant next to the difference between an answer you can use and an answer you have to rewrite. Speed matters when the assistant sits inside a customer-facing loop.
Step four: check the things you cannot see in the output
Three questions decide whether a good tool is a safe tool for your context. First, where does your input go, and is it used to improve the vendor's models? Consumer and business tiers of the same product frequently answer this differently, and the business answer is usually the one you want if you handle client material. Second, can you export your conversations, prompts, and generated content in a usable format? Anything you cannot export is content you are renting. Third, does your organisation, client contract, or regulator place restrictions on where data is processed?
These are not exciting questions, and they are the ones that turn a good pilot into a problem six months later. Vendors publish the answers, and they change — check them at decision time rather than trusting a summary written by anyone, including us.
Step five: decide on total effort, not headline price
Pricing for assistants moves often, so treat any number you read anywhere — including here — as needing confirmation on the vendor's own page. What is more stable is the shape of the cost. A per-seat subscription is predictable and easy to approve. Usage-based pricing is cheaper when work is occasional and harder to forecast when it becomes central. A free tier is genuinely useful for evaluation and is rarely a sound base for work you depend on, because free tiers are the first thing that changes.
Add the cost you will not see on an invoice: the time spent moving material into the tool, the time spent verifying output, and the time spent training people to use it well. A slightly more expensive assistant that lives inside the software your team already uses often wins on this line alone.
A note on switching
The most reassuring thing about this category is how low the switching cost has become. Prompts are portable. Habits transfer. If you choose reasonably now and revisit in six months, you have lost very little. That argues for deciding quickly on evidence from your own material, rather than waiting for a clearer winner that the market shows no sign of producing.
You can browse the assistants covered on this site in the HOMERA-X directory. Each profile links to the vendor's own website and pricing page so you can verify current terms directly.
Tools mentioned
Each profile links to the vendor's own website and pricing page.
Tags: assistants · selection · workflow
Back to AI Guides

