A citation can make an AI-assisted answer look safer than an uncited answer. It can also create a false sense of completion. For a contentious lawyer, the real question is not simply whether a footnote appears. It is whether the authority or evidence exists, whether the cited passage is genuine, whether that passage supports the proposition actually advanced, and whether the proposition remains defensible when the rest of the relevant material is considered.

That distinction matters because litigation is full of genuine documents that do not prove what a party says they prove. A real email may concern the wrong meeting. A real judgment may state a principle subject to an exception omitted from the summary. A genuine witness statement may record recollection that is qualified by contemporaneous documents. Verification therefore has to operate at the level of the proposition, not merely at the level of the hyperlink.

1. Four different questions hide behind the word “citation”

A disciplined review can separate four stages. First: does the cited source exist? Second: does it contain the passage or proposition attributed to it? Third: does that material actually support the proposition being advanced? Fourth: is the proposition complete when considered against the rest of the relevant record?

The first two stages catch obvious fabrication and miscitation. The third catches a subtler problem: a real source can be cited for a proposition it does not establish. The fourth is the most litigation-specific. A proposition may be accurately supported by one document and still be misleading because another document qualifies it, contradicts it or changes its significance.

2. The courts are concerned with accuracy and responsibility, not merely fabricated authorities

In R (Ayinde) v London Borough of Haringey, the Divisional Court stressed the need to check AI-assisted legal research against authoritative sources before relying on it. The lesson is often reduced to “do not cite fake cases”. That is too narrow. The professional obligation is to check the work that is being relied upon.

Hancox v Sutherland makes the wider point especially clearly. The Employment Appeal Tribunal dealt with a 300-page skeleton argument created using ChatGPT. The court’s concerns included not only accuracy but also excessive length, relevance, procedural compliance and the burden placed on the tribunal and the other party. The judgment states that litigants using generative AI must take personal responsibility for documents submitted to the EAT and check them as thoroughly as they reasonably can.

3. Citation quality and answer correctness are separate dimensions

Research on citation-generating language models supports the same conceptual distinction. The ALCE benchmark evaluates fluency, correctness and citation quality separately. That is useful for lawyers because it resists the assumption that an answer becomes correct merely because it is accompanied by sources.

A 2024 study of then-current proprietary legal research tools similarly found that retrieval-augmented systems materially reduced hallucination compared with a general-purpose comparator but did not eliminate it. Those percentages are historical observations about the products and versions tested in 2024; they should not be presented as measurements of their 2026 performance. The enduring point is methodological: retrieval and citation can improve verification without turning the output into an unquestionable conclusion.

4. Source support should be tested at proposition level

Consider the proposition: “The defendant knew about the defect before completion.” A source-linked review should be able to ask who is said to have known, when knowledge is said to have arisen, which document records it, whether that person authored or received the document, whether the language records actual knowledge or only suspicion, and what material points the other way.

That is different from attaching three documents to the sentence. The useful unit of review is the relationship between a proposition and the material said to support, qualify or contradict it.

5. Verification also requires scope awareness

An answer can be faithful to every passage retrieved and still be incomplete because the retrieval step missed the decisive document, selected the wrong version, omitted an attachment or searched a corpus that did not contain the relevant material. This is why “the system found no document” should not become “no document exists”. Retrieval adequacy and corpus completeness are separate questions from generation accuracy.

Contentious practice already provides a useful analogue. Practice Direction 57AD treats adverse material as important even where it damages the disclosing party’s own account. It also recognises that search design matters. Legal AI is not governed by PD57AD merely because it uses retrieval, but the procedural framework illustrates a core litigation discipline: an evidential conclusion depends partly on what was searched and what was available to be searched.

6. A practical four-stage verification protocol

  1. Existence: locate the actual authority or matter document.
  2. Passage: confirm the cited wording, paragraph, page or message.
  3. Support: decide whether that material really supports the proposition, and whether the proposition goes beyond it.
  4. Completeness: look for material that qualifies, contradicts or materially changes the conclusion.

The fourth step is what turns citation checking into contentious-law analysis. It is also where source-linked matter work becomes more useful than a citation badge attached to a fluent answer.

7. Legal authority and matter evidence require different checking

The mechanics of verification differ depending on what kind of proposition is being tested. A legal proposition may require checking legislation, a Practice Direction or the ratio and context of a judgment. A factual proposition may require checking emails, contracts, notes, witness evidence, metadata or other matter material. A mixed proposition may require both.

This distinction prevents a common category error. A case can establish the legal test without proving that the test is satisfied on the facts. Conversely, a document can strongly support a factual event without answering what legal consequence follows. Source-linked work should therefore make the source type visible rather than treating every citation as equivalent.

8. Partial support needs to remain partial

AI-generated sentences often combine several propositions. A source may support only one part. For example, “The board approved the transaction on 10 June and informed the claimant the same day” contains at least two factual propositions. A board minute may support the first but say nothing about communication to the claimant.

Verification at sentence level can therefore be too coarse. Where the point matters, the reviewer should decompose the sentence and test each material proposition separately. This is particularly important where timing, knowledge, causation or state of mind is disputed.

9. Team review depends on reproducibility

In a substantial matter, the person who generated an AI-assisted finding may not be the person who later relies on it. A partner, counsel or another associate may need to understand what was searched and why the proposition was accepted. A good verification trail should therefore survive handover.

That does not require recording every prompt forever. It does require enough matter context to reconstruct the important decision: proposition, source, relevant passage, contrary material, uncertainty and professional outcome. This is closer to litigation knowledge management than to preserving a chat transcript.

10. Verification is a cost problem as well as a risk problem

If checking an AI result requires manually reopening dozens of documents and reconstructing how the answer was reached, the efficiency benefit may disappear. Source-linked design should reduce the cost of professional scepticism by making the supporting and qualifying material immediately inspectable.

The commercial question for legal AI is therefore not simply whether the model can answer quickly. It is whether the whole cycle—question, retrieval, checking, challenge and professional decision—can be performed more effectively without weakening the standard of review.

Conclusion

A citation is an inspection route, not a certificate of truth. The professionally useful question is whether a lawyer can move from the proposition to the source, test the relationship between them, identify contrary material and decide what the evidence and law actually justify.

Related LegalRAG Pro Insights

Professional context. This article discusses legal-technology workflow and professional-risk questions. It is not legal advice and should not be treated as a substitute for checking the current procedural, regulatory and factual position in a particular matter.