Choose an AI by starting with the work to be done, the data you permit it to receive, and the way you will check its answer. A general assistant is a test direction for a broad but bounded request. Sourced research fits a need for traceability. A workflow integration is worth evaluating once the task and its checks are defined. If the output, data boundary, or validation remains unclear, do not choose a tool yet. Clarify the task. The right AI is not a universal winner. It is a candidate that passes a test built around your case. Give every candidate the same input, request the same deliverable, and apply criteria written before the trial. Check current pricing, availability, terms, and data practices on official provider pages when you run it.
My operator’s verdict: I would not choose an AI from a general leaderboard. That is my editorial judgement based on this journal’s method, not a claim based on private product testing. I would retain the candidate whose result is repeatable on the stated task, verifiable through the agreed checks, and compatible with the data and risk boundary set before the trial.
What result are you actually asking for?
“Which AI should I choose?” comes too early when the desired result is vague. “Help me work” supplies no decision criterion. Preparing a first structure, finding material whose sources must be checked, and placing a defined activity inside a business workflow are different requests.
The first decision is therefore not a brand. It is the shape of the expected output. Before opening a tool, write down:
- what the candidate receives;
- what it must return;
- whether each usable claim must be traceable to a source;
- what a human will check before using the output;
- which condition will stop the trial.
This separates a pleasant exchange from a usable deliverable. It also gives every candidate the same ground. If the prompt, expected output, or validation changes between trials, the comparison no longer answers the original decision.
If you cannot describe the deliverable, the current verdict is “define the task”, not “choose an AI”.
| Dominant need | What to observe | Test direction |
|---|---|---|
| Broad request or exploratory work | Usefulness of the deliverable for the stated task | General assistant |
| Answer that must lead back to evidence | Source traceability and source checking | Sourced research |
| Defined work in an operating context | Fit with the allowed data and planned check | Workflow integration |
| Unclear output, data, or check | Ability to formulate a test | Define the task first |
The benefits and limits of AI depend on the context in which it is used. A related guide on AI’s benefits and drawbacks is available in French.
What data are you willing to share?
A candidate may fit the deliverable while remaining outside the boundary set for its input. Before the trial, I separate what may enter the tool from what must remain out. If that line is uncertain, I suspend the choice.
I use three questions:
- What information will the tool receive, exactly?
- Which parts have I decided may be transmitted?
- Can I run the same test without data whose status is unresolved?
Current data practices must be checked in the documentation of the provider under consideration. The Claude privacy help centre and the Gemini Apps privacy hub are stable official places to check the current practices of those providers. These links do not support a broad privacy claim or a comparison between products. The Mistral conversation documentation is an official documentation surface, not evidence of comparative quality.
Terms, availability, and prices also need to be checked on each candidate’s official page at the time of the trial. I do not reproduce them here because the evidence packet does not establish a current product-by-product statement.
“I do not know which data I can share” is a useful diagnostic result. Return to the task boundary before selecting a tool.
This filter also comes before any decision about an agent. The guide to creating an AI agent is available in French and begins from the same framing need.
Which working mode does the task require?
A sourced answer is evaluated by whether a reader can reach the cited material and check it. An extended deliverable is evaluated against the full output described before the trial. A workflow integration is evaluated inside the stated task, allowed data, and planned verification. The expected working mode determines the grid.
That leaves four useful directions:
- Define the task when the result, data boundary, or check is uncertain.
- Test a general assistant for a versatile, bounded request.
- Test sourced research when traceability forms part of the output.
- Evaluate a workflow integration when the work and its validation are already framed.
Cost can enter the decision after this framing, but it does not define the need. The separate guide on what an AI agent costs is in French.
How can you copy a comparison grid without fake scores?
This is the sheet I would use. Give each candidate its own row or copy of the table. Do not add a total score or hidden weighting. Fill the cells from what the trial produced, not from a product reputation.
| Field to copy | What to record |
|---|---|
| Task | The precise work given to every candidate |
| Fixed input | The same text, document, or permitted information |
| Expected output | The deliverable format defined before the trial |
| Source traceability | Required sources and planned check, when needed |
| Data boundary | What may enter the tool and what remains excluded |
| Human check | Who checks the result and against which criteria |
| Stop condition | The defect or uncertainty that ends the trial |
| Observed result | A factual description of what the candidate returned |
To use it:
- duplicate the sheet for every candidate;
- leave the fixed input, expected output, and human check unchanged;
- describe observations without turning them into a universal ranking;
- stop when a written stop condition is met;
- retain the input and verdict so the decision can be explained.
The NIST AI Risk Management Framework is intended for voluntary use, helps manage AI risk, and incorporates trustworthiness considerations. I use it as a risk-management reference, not as an endorsement of any candidate.
What does reliable AI mean for your case?
“Reliable” does not mean universally superior in this guide. I define it locally: the result is repeatable on the stated task, can be verified through the planned grid, and stays within the risk boundary chosen for the trial. That conclusion applies only to the examined case.
The check can cover:
- whether the deliverable matches the expected output;
- whether required sources can be reached and checked;
- whether the defined data boundary was respected;
- whether a stop condition appeared;
- whether the observed result is recorded without extrapolation.
An impressive result on another task does not decide this one. A displayed citation is not enough either. When a claim requires a source, the human check is to open that source and confirm that it supports the claim being used.
When should a human remain in the loop?
I keep a human in the loop whenever the deliverable needs checking before use. Their job is not to repair a vague request. It is to apply the control written before the trial: inspect sources, inspect the deliverable, enforce the data boundary, and observe the stop condition.
Return to framing when:
- you cannot state which result you would accept;
- you are uncertain about the data to transmit;
- you cannot say how the answer will be checked.
In each case, adding a tool does not settle the uncertainty. The decision begins with the work and continues with the candidate that can pass its test. For a wider perspective, the journal’s guide to the future of AI is available in French.
What are the short answers to common questions?
Which AI is best, or the most reliable?
I do not name a universal winner. For a stated task, I call a candidate reliable when its result is repeatable with the fixed input, verifiable through the grid, and contained within the chosen risk boundary. The verdict remains local to that trial.
Which free AI should I choose, and what does it currently cost?
Check current availability, any free access, limits, and price on the provider’s official page when you run the trial. This guide states no price. A candidate presented there as free still needs to pass the same task, data, and verification test.
Should I choose Claude or ChatGPT?
The names alone do not decide it. Give both candidates the same task, input, expected output, and check. Consult their current official pages for terms, pricing, and availability, then keep the verdict limited to your case.
How should I handle data and verify an answer?
Set the permitted data boundary before the trial and consult the provider’s current official documentation. Then write down the human check, the sources to open when sources are required, and the condition that ends the trial.
You now have four doors: define the task, test a general assistant, test sourced research, or evaluate a workflow integration. The next step follows the stated need, not a leaderboard.