Questions to Ask When Evaluating Legal AI

Short answer: What decides a legal AI evaluation is not the capability list but the answer to five questions: does the output cite sources, does the cited source actually support the claim, where is the data processed, are your files used for model training, and can the system say it does not know. If any one of these is unanswered, the remaining features carry little practical value.
AI is no longer a differentiator in legal software; nearly every vendor offers an assistant. The real question is therefore not whether the system has AI, but whether a lawyer can defend the output it produces.
The questions below fall into five areas. Each closes with a way to test that area during the demo.
1. Accuracy and source attribution
Does every claim link to a source? Each decision, provision and date in an answer should resolve to a clickable source. A system that shows no sources cannot be verified even when it is right, which makes it unusable for legal work.
Does the source actually support the claim? This is a separate question. A system can link to a real decision and assert something that decision does not say. Test the accuracy of the citation, not its presence.
Can the system say it does not know? A system that states a question is out of scope is more trustworthy than one that answers everything with confidence.
How to test it: Ask about something you already know and open every citation. Then ask about a regulation that does not exist. If the system confirms it and starts explaining, its tendency to fabricate is high.
2. Data processing and confidentiality
Where are uploaded files processed? If the model runs abroad, a cross-border transfer is involved and needs a documented basis. Ask in writing which model providers the vendor uses as sub-processors.
Are files used for model training? Look for an explicit clause. It must bind not only the vendor but also the model providers behind it.
Can client data be anonymised first? Professional secrecy cannot be delegated to a vendor. Being able to mask names, identifiers and matter numbers before the assistant sees them makes that duty far easier to manage.
How to test it: Upload a file and ask the vendor to walk you through every system it passes. If they cannot draw a clear path, the data journey is unclear to them as well.
3. Coverage and currency
Which sources are indexed? Supreme court, administrative court, appellate courts, constitutional court, legislation: which of these are in scope? Any accuracy claim made without a coverage list is unmeasurable.
How often is the data refreshed? Ask how many days a newly published decision takes to appear. Out-of-date case law is more dangerous than wrong case law, because it looks right.
Is the status of a decision shown? Does the system flag whether a decision was later overturned or superseded? Without that, the research is incomplete.
How to test it: Search for a recent decision whose outcome you know. If it is missing, the refresh cycle is slow. If present, check whether the system indicates its current status.
4. Fit with the workflow
Where does the output go? Can the generated text move into the matter file or a filing, or does it have to be copied and pasted? The second option gives back the time the tool saved.
Does it see matter context? Does the assistant know the parties, stage and attachments of the matter you are working on, or does every question start from zero? A context-free assistant is no different from a general-purpose chat tool.
Can output be shared within the team? Attaching a research result to the matter and handing it to a colleague matters more in corporate use than individual speed.
How to test it: Run one real matter end to end: research, draft, save to the matter, hand to a colleague. Note every step that requires copy and paste.
5. Accountability and audit
Is there a record of who generated what and when? When an error surfaces later, you need to see when the output was produced and from which prompt.
Can firm policy be enforced in the system? Can you restrict which team members may use the assistant on which types of matter?
What does the vendor commit to for faulty output? Most vendors leave responsibility entirely with the user. That may be reasonable, but you should sign knowing what the contract says.
How to test it: Ask to see the list of past queries. If no such screen exists, auditing is a promised feature rather than a present one.
Acceptance thresholds
Put a threshold on each area to make the assessment concrete. The thresholds below are a minimum for corporate use.
| Area | Minimum threshold | How to measure |
|---|---|---|
| Source attribution | Every legal claim links to a clickable source | Pick ten claims and open each citation |
| Citation accuracy | At least nine of ten sources support the claim | Read the same ten citations as content |
| Out-of-scope behaviour | Does not confirm a non-existent regulation | Ask about a fictional provision |
| Data processing | Processing location and sub-processors supplied in writing | Request the contract annex |
| Model training | Clause excluding training is in the contract | Ask for the clause number |
| Currency | Time to index a new decision is committed | Search for a decision from the last month |
| Workflow | Research to matter file needs no copy and paste | Run one flow end to end |
| Audit | Past queries can be viewed | Ask to see the screen |
Four common mistakes
- Passing the demo with the vendor's own question. Their sample questions show where the system is strong. Bring your own hard question.
- Treating the presence of a citation as accuracy. Producing a link is easy; producing a link that supports the claim is not. Test the difference.
- Skipping confidentiality without legal input. Processing location and model training are legal questions, not technical ones, and that view belongs at the decision table.
- Deciding on one impressive output. These systems produce variable results. Ask the same question on different days.
Frequently asked questions
What is the single most important question when evaluating legal AI?
Whether every legal claim in the output links to a clickable source. A system that shows no sources cannot be verified even when it is correct, which makes it unusable for legal work. Separately, the accuracy of the citation must be tested: a system can link to a real decision and assert something that decision does not say.
Who is responsible if the AI produces a fabricated citation?
Most vendor contracts leave responsibility with the user, and the lawyer is accountable for a citation in a filing submitted to a court. For that reason, whether the system shows sources and whether its output can be checked matters more than the liability clause in the contract.
Is uploading client files to an AI assistant a data protection problem?
Not inherently, but it needs a documented basis. If the model runs abroad, a cross-border transfer is involved. Professional secrecy also cannot be delegated to a vendor, so the ability to anonymise before upload and an explicit contractual clause excluding model training both matter.
How do we understand the coverage of a legal AI system?
Ask for a written list of indexed sources: supreme court, administrative court, appellate and constitutional courts, and legislation. Accuracy claims made without a coverage list are unmeasurable. Also ask how many days a newly published decision takes to appear in the system.
What is the difference between a demo and a trial?
A demo is a presentation the vendor controls, with questions and examples chosen to show the system at its best. A trial is a test you control, using your own matters, your own hard questions and your own workflow. The decision should rest on the trial.
How many examples are enough to measure accuracy?
One is not enough, because these systems produce variable results. A practical threshold is to take ten claims and both open and read their citations as content. Repeating the same question on different days also measures consistency.



