Start with a job you can check
Write down one result you need: a draft you can edit, a summary of a specific document, an explanation you can verify, or a repeated task you can supervise. 'Use more AI' is too broad to compare tools. A fluent answer is useful only if it meets your actual requirements.
For example, compare two tools on the same anonymized meeting notes. Ask each to extract decisions, owners and due dates, then check every item against the notes. Record missing details and invented commitments. This is a proposed test, not a report of vendor results.
From the notes below, make a table of decisions, named owners and deadlines. Use 'not stated' for missing fields. Quote the supporting sentence for each row. Do not infer an owner or deadline.
Choose how much setup you want
A ready-to-use app lets you work in a browser or installed interface. An automation connects accounts and repeats steps. A framework or API is a building block that usually needs someone to implement it. These can overlap within one vendor; choose the specific offering rather than the brand name.
If you only need to draft or explain something occasionally, start with an app you can use directly. If the same task moves information between accounts every week, compare a workflow. Developer tools make sense when you need a custom application, permissions or deployment that an existing app cannot provide.
In ASE, the type label tells you what sort of product is listed. 'MCP server' means a connector for a compatible AI client, not a standalone chat app. 'Repository language' describes the code, not the languages a person can speak to the product.
Four ready-to-use apps to put through your own test
You do not need an agent framework to begin. These are documentation-based examples, not a ranking or a report of our product tests. Choose two that support your task and are available to your account. Product documentation checked 8 October 2026; features, usage limits and data controls can differ by plan and location.
ChatGPT: OpenAI documents drafting, rewriting, summarizing and explanations, with additional tools depending on plan and settings. A useful first test is the same short draft plus a fact-checking pass. If the question needs current evidence, check whether the relevant search or research tool was actually used.
Claude: Anthropic documents a conversational assistant available on web, desktop and mobile for tasks including analysis and creative writing. Try the same source notes and writing brief, then inspect omissions and changed commitments. Check supported locations and account requirements rather than assuming universal availability.
Gemini: Google documents its app and connections to other services for tasks such as finding information or working with email and calendars. Start with a plain-text task before connecting personal accounts. Available connections and actions vary by account, device, country and other conditions; read the controls for your setup.
Perplexity: its search documentation describes research answers with sources. Test a question whose evidence you can independently inspect, including one conflicting or outdated source. A sourced answer still needs verification, and the specific search mode and limits depend on the offering you select.
Compare two candidates with the same inputs
Choose two options that can produce the same kind of output. Give both the same source material, instructions and constraints. Keep a copy of the output, the plan used, the date and the time you spend correcting it. A product may work well for a short email and poorly for a document with tables.
Include one ordinary case and one awkward case: missing information, a conflicting date, an unreadable table or a request outside the tool's knowledge. An honest 'I cannot tell from this source' can be a better result than a confident invention.
- Accuracy: check facts and omissions against the original.
- Usefulness: can you use the output after reasonable edits?
- Effort: record setup, checking and correction time.
- Control: can you preview changes, undo them and export your work?
- Cost: record the plan and any usage or connected-service charges.
Check what the tool can read and change
Start with public, fictional or anonymized material. Before uploading private work, read the provider's current data controls for the actual plan. Check retention, model-training choices, connected apps, sharing and deletion. A setting available in a business plan may differ from a personal account.
For an automation, distinguish permission to read from permission to send, delete, purchase or publish. Test a draft-only workflow first. Confirm the actual account state after a run; an agent saying 'done' does not establish that the correct change happened.
Running open-source software yourself does not automatically keep all data on your device. It may still send prompts, files or logs to a model API or hosted service. Trace those dependencies before choosing by a privacy label.
Make the test fit your country and language
Check service availability and account requirements in your country. Test the language, accent, script and terminology you actually use; an English interface does not establish good results in every language. For a voice tool, use the channel and noise conditions your callers will encounter.
Compare prices in the quoted currency, including taxes, exchange costs and separate model or API charges. Write dates unambiguously, such as 8 October 2026, and specify a time zone when scheduling. A tool that assumes a different date format can produce a plausible but wrong result.
Use British English. Keep all monetary amounts in their original currencies. Write dates as day, month name and year. If a date or time zone is ambiguous, ask before changing it.
Measure the cost of a useful result
A free tier can help you test, but limits, uploads, exports or automations may differ from a paid plan. Software licensing, subscription fees and model usage are separate costs. Check the current provider page before buying; ASE's pricing category is a starting point, not a quote.
Example arithmetic: a hypothetical US$20 monthly plan used for 40 acceptable tasks costs US$20 ÷ 40 = US$0.50 per acceptable task, before tax and your time. If 10 of those outputs need 6 minutes of correction each, that adds 10 × 6 = 60 minutes of work. These are illustrative inputs, not any vendor's price or measured performance.
Choose the option that meets your minimum quality and control requirements at a cost you can sustain. Keep your existing method if neither candidate improves the complete task. A new subscription is not the only successful outcome of an evaluation.
Common questions
- Do I need to code to use AI?
- No. Many products offer a browser or app interface. APIs, frameworks and some connectors are for people building or integrating software. Start from the setup required for the specific product.
- Does a high GitHub star count mean a better AI tool?
- No. It measures repository interest. It does not establish accuracy, active users, support, privacy or suitability for your task.
- Can one AI tool do everything?
- Some products cover many tasks, but their capabilities, limits and quality can differ by task. Test the work you need instead of assuming a broad feature list guarantees a useful result.
AI-assisted drafting and source review by Agent Search Engine. Fictional examples and proposed tests are labeled. Read our editorial method or report a correction.
More practical guides
Everyday AI / Decision guide
How to research and learn with AI without losing the sources
A practical AI research workflow: define a question, collect sources, verify citations and separate evidence from inference. Includes a worked source-checking exercise.
Read the guideEveryday AI / Decision guide
How to write with AI and keep control of the final draft
Use AI for emails, summaries and first drafts with a clear brief, a fact ledger and a review checklist. Includes practical prompts and a fictional before-and-after exercise.
Read the guideVoice / Decision guide
Choose and evaluate a voice agent stack
Vapi, Retell, LiveKit and Pipecat: compare architectures, test twelve difficult call scenarios, and model the complete cost.
Read the guide

