Questions to Ask When Evaluating Legal AI

Questions to Ask When Evaluating Legal AI
Gökçen Beyazoğlu

Gökçen Beyazoğlu LL.B.

Chief Product Officer · Attornaid

Short answer: What decides a legal AI evaluation is not the capability list but the answer to five questions: does the output cite sources, does the cited source actually support the claim, where is the data processed, are your files used for model training, and can the system say it does not know. If any one of these is unanswered, the remaining features carry little practical value.

AI is no longer a differentiator in legal software; nearly every vendor offers an assistant. The real question is therefore not whether the system has AI, but whether a lawyer can defend the output it produces.

The questions below fall into five areas. Each closes with a way to test that area during the demo.

1. Accuracy and source attribution

Does every claim link to a source? Each decision, provision and date in an answer should resolve to a clickable source. A system that shows no sources cannot be verified even when it is right, which makes it unusable for legal work.

Does the source actually support the claim? This is a separate question. A system can link to a real decision and assert something that decision does not say. Test the accuracy of the citation, not its presence.

Can the system say it does not know? A system that states a question is out of scope is more trustworthy than one that answers everything with confidence.

How to test it: Ask about something you already know and open every citation. Then ask about a regulation that does not exist. If the system confirms it and starts explaining, its tendency to fabricate is high.

2. Data processing and confidentiality

Where are uploaded files processed? If the model runs abroad, a cross-border transfer is involved and needs a documented basis. Ask in writing which model providers the vendor uses as sub-processors.

Are files used for model training? Look for an explicit clause. It must bind not only the vendor but also the model providers behind it.

Can client data be anonymised first? Professional secrecy cannot be delegated to a vendor. Being able to mask names, identifiers and matter numbers before the assistant sees them makes that duty far easier to manage.

How to test it: Upload a file and ask the vendor to walk you through every system it passes. If they cannot draw a clear path, the data journey is unclear to them as well.

3. Coverage and currency

Which sources are indexed? Supreme court, administrative court, appellate courts, constitutional court, legislation: which of these are in scope? Any accuracy claim made without a coverage list is unmeasurable.

How often is the data refreshed? Ask how many days a newly published decision takes to appear. Out-of-date case law is more dangerous than wrong case law, because it looks right.

Is the status of a decision shown? Does the system flag whether a decision was later overturned or superseded? Without that, the research is incomplete.

How to test it: Search for a recent decision whose outcome you know. If it is missing, the refresh cycle is slow. If present, check whether the system indicates its current status.

4. Fit with the workflow

Where does the output go? Can the generated text move into the matter file or a filing, or does it have to be copied and pasted? The second option gives back the time the tool saved.

Does it see matter context? Does the assistant know the parties, stage and attachments of the matter you are working on, or does every question start from zero? A context-free assistant is no different from a general-purpose chat tool.

Can output be shared within the team? Attaching a research result to the matter and handing it to a colleague matters more in corporate use than individual speed.

How to test it: Run one real matter end to end: research, draft, save to the matter, hand to a colleague. Note every step that requires copy and paste.

5. Accountability and audit

Is there a record of who generated what and when? When an error surfaces later, you need to see when the output was produced and from which prompt.

Can firm policy be enforced in the system? Can you restrict which team members may use the assistant on which types of matter?

What does the vendor commit to for faulty output? Most vendors leave responsibility entirely with the user. That may be reasonable, but you should sign knowing what the contract says.

How to test it: Ask to see the list of past queries. If no such screen exists, auditing is a promised feature rather than a present one.

Acceptance thresholds

Put a threshold on each area to make the assessment concrete. The thresholds below are a minimum for corporate use.

AreaMinimum thresholdHow to measure
Source attributionEvery legal claim links to a clickable sourcePick ten claims and open each citation
Citation accuracyAt least nine of ten sources support the claimRead the same ten citations as content
Out-of-scope behaviourDoes not confirm a non-existent regulationAsk about a fictional provision
Data processingProcessing location and sub-processors supplied in writingRequest the contract annex
Model trainingClause excluding training is in the contractAsk for the clause number
CurrencyTime to index a new decision is committedSearch for a decision from the last month
WorkflowResearch to matter file needs no copy and pasteRun one flow end to end
AuditPast queries can be viewedAsk to see the screen

Four common mistakes

  • Passing the demo with the vendor's own question. Their sample questions show where the system is strong. Bring your own hard question.
  • Treating the presence of a citation as accuracy. Producing a link is easy; producing a link that supports the claim is not. Test the difference.
  • Skipping confidentiality without legal input. Processing location and model training are legal questions, not technical ones, and that view belongs at the decision table.
  • Deciding on one impressive output. These systems produce variable results. Ask the same question on different days.

Frequently asked questions

What is the single most important question when evaluating legal AI?

Whether every legal claim in the output links to a clickable source. A system that shows no sources cannot be verified even when it is correct, which makes it unusable for legal work. Separately, the accuracy of the citation must be tested: a system can link to a real decision and assert something that decision does not say.

Who is responsible if the AI produces a fabricated citation?

Most vendor contracts leave responsibility with the user, and the lawyer is accountable for a citation in a filing submitted to a court. For that reason, whether the system shows sources and whether its output can be checked matters more than the liability clause in the contract.

Is uploading client files to an AI assistant a data protection problem?

Not inherently, but it needs a documented basis. If the model runs abroad, a cross-border transfer is involved. Professional secrecy also cannot be delegated to a vendor, so the ability to anonymise before upload and an explicit contractual clause excluding model training both matter.

How do we understand the coverage of a legal AI system?

Ask for a written list of indexed sources: supreme court, administrative court, appellate and constitutional courts, and legislation. Accuracy claims made without a coverage list are unmeasurable. Also ask how many days a newly published decision takes to appear in the system.

What is the difference between a demo and a trial?

A demo is a presentation the vendor controls, with questions and examples chosen to show the system at its best. A trial is a test you control, using your own matters, your own hard questions and your own workflow. The decision should rest on the trial.

How many examples are enough to measure accuracy?

One is not enough, because these systems produce variable results. A practical threshold is to take ten claims and both open and read their citations as content. Repeating the same question on different days also measures consistency.