Skip to content

A gradual release of Avioverse begins in October 2026. Request early access →

AI Hallucinations and EASA Rules: How to Check an AI Answer Against the Regulation

Can you trust an AI answer on EASA rules? What the studies measured, and a five-step check of any AI answer against the rule text, worked on 145.A.40.

Dionysis Kefalas11 min read

On this page

Day two of an internal audit at ExampleMRO. In the tool store, torque wrench TW-EXAMPLE-07 carries a calibration label that ran out three weeks ago. You need the requirement for the finding, so you ask an AI assistant. The reply takes four seconds:

Under 145.A.40(b), all tools and test equipment must be calibrated at least every 12 months against a national standard, and calibration records must be kept by the organisation [1].

(Illustrative reply, written for this article.)

A rule number, a sub-paragraph, an interval and a citation. It reads like the rule. It is not what the rule says.

So can you trust ChatGPT, or any AI assistant, on EASA regulations? Use the answer as a pointer to where the rule is, never as the rule. Published studies keep finding cited AI answers that their own sources do not back up. A citation can be real and still sit on a sentence the passage does not say. The check takes a few minutes: open the cited passage, confirm the revision, check whether it is IR, AMC or GM, find the conditions and exceptions, and compare the sentence with the passage word by word.

Below: what the research measured, what EASA has said, the five steps worked on this reply, and how to do the same in Avioverse.

What the research measured

We found no peer-reviewed or institutional study that tests AI answers against EASA rule text, as of 25 September 2026. The nearest evidence comes from law, medicine and news, where researchers checked whether cited sources support what the answer says. Each figure belongs to its own test. None is a general error rate for AI.

StudyWhat was testedWhat it found
Wallat et al., ICTIR 2025 (first posted December 2024)One citation-producing model on one question set. Researchers copied short cited phrases from its answers, usually about four words, into other documents in its context.With the phrase planted in a relevant but uncited document, the model cited that document for the same statement in 57% of cases (273 of 476). The planted text matched the words, not the claim, and the model had given the same answer without it.
Tow Center for Digital Journalism, Columbia Journalism Review, 6 March 2025Eight AI search tools, 1,600 queries: given a verbatim news excerpt, name the headline, publisher, date and URL.Wrong on more than 60% of queries overall, from 37% to 94% by tool. Every tool except Copilot, which declined more questions than it answered, was "consistently more likely to provide an incorrect answer than to acknowledge limitations".
Wu et al., Nature Communications, April 2025Seven models, 800 medical questions, about 58,000 statement–source pairs.50% to 90% of responses were "not fully supported, and sometimes contradicted, by the sources they cite". One model running with web search still left about 30% of individual statements unsupported.
EBU and BBC, 21 October 2025Journalists from 22 public-service media organisations in 18 countries rated 2,709 answers from the free consumer versions of ChatGPT, Copilot, Gemini and Perplexity.45% had at least one significant issue. 31% had significant sourcing problems, which include information the cited source does not support.
Stanford RegLab (Magesh et al.), Journal of Empirical Legal Studies, 2025202 preregistered legal queries, run from March to May 2024 on commercial legal research tools and a general-purpose model.The legal tools hallucinated on 17% to 33% of queries, the general-purpose model on 43%. A real citation that does not support the claim counted as a hallucination.

The legal tools in the Stanford study were 2024 versions answering questions on US law, so the rates are not today's rates for any product. What carries over is the definition. The authors call an answer "misgrounded" when it cites a source that does not support the claim, and they count it as a hallucination. Wallat's result points the same way from the other side: a citation can match the words of a statement without containing the claim, or being where the answer came from. Either way, a citation tells you where to look. It does not tell you the sentence is right.

The torque-wrench reply is misgrounded. 145.A.40(b) exists. It does not say 12 months.

Courts have already drawn the professional line. In Ayinde v Haringey [2025] EWHC 1383 (Admin), handed down on 6 June 2025, the High Court of England and Wales warned that such tools "may cite sources that do not exist" and "may purport to quote passages from a genuine source that do not appear in that source". It set out a professional duty "to check the accuracy of such research by reference to authoritative sources". In the second case heard with it, 18 of 45 citations were to cases that did not exist.

Aviation has its own case. In Moffatt v Air Canada, 2024 BCCRT 149 (14 February 2024), a British Columbia tribunal held the airline liable for negligent misrepresentation after its website chatbot misstated the airline's bereavement fare policy. It called the suggestion that the chatbot was, in effect, a separate legal entity "a remarkable submission": the airline is "responsible for all the information on its website". Nor did the customer have to check the chatbot against the airline's own bereavement travel page. The organisation owned what its AI said.

What EASA has said about AI answers

We found no adopted EASA rule for AI assistants as of 25 September 2026. What exists is guidance and a proposal:

  • AI Concept Paper Issue 2 (March 2024) is guidance for Level 1 and Level 2 machine-learning applications. Level 1 is "assistance to human"; Level 2 is "human-AI teaming".
  • NPA 2025-07 (10 November 2025) is EASA's first regulatory proposal on AI trustworthiness. It is limited to Levels 1 and 2 and says that "generative AI tools and general-purpose models are not yet fully covered". Its consultation has closed; we found no decision adopting it.
  • Concept Paper proposed Issue 03 (3 June 2026) brings large language models into scope and names the failure: "Emergent behaviour and unexpected outputs (with a subtype also referred to as ‘hallucination’): LLMs can exhibit emergent behaviour, where the model generates unexpected or unconventional outputs that are not necessarily aligned with the intended goal(s) or task(s)." For Level 1 AI, "decisions are taken by the end user based on support by the AI-based system."

These documents target safety-related AI used in approved products or by approved organisations. None of them approves an assistant or turns its answer into a source you can cite in a finding. EASA AI levels: assistance vs decision-making covers the levels in detail.

How to check an AI answer against the regulation in five steps

These steps work with any tool. Each one below is worked on the torque-wrench reply.

1. Open the cited passage, not a summary

Follow the citation to the paragraph itself: the Easy Access Rules, the Official Journal, or a reader that shows the original wording. A snippet in the chat or the model's paraphrase is not the passage. If the citation points nowhere, or you cannot find the paragraph, the answer has no basis yet.

Here is the paragraph the reply cites:

145.A.40(b) · Easy Access Rules for Continuing Airworthiness, 2025-09-02 revision

(b) The organisation shall ensure that all tools, equipment and particularly test equipment, as appropriate, are controlled and calibrated according to an officially recognised standard at a frequency to ensure serviceability and accuracy. Records of such calibrations and traceability to the standard used shall be kept by the organisation.

The full point is on the 145.A.40 rule page. The citation is real, which is what makes the reply dangerous. A real reference passes a quick glance.

2. Check it is the revision in force

The model may have learnt an older text, and the copy you have open may be older than the date that matters. Compare the revision date of the document with the date the finding refers to. Watch for points printed twice: in the Air Ops rules, AMC1 ORO.GEN.200(a)(3) carries blocks tagged applicable until 31 December 2027 and applicable from 1 January 2028. An answer that quotes the 2028 block today is quoting text not yet in force.

For the example, 145.A.40(b) above comes from the 2 September 2025 revision and carries no such tags. Before the finding goes out, confirm on EASA's Easy Access Rules page that no later revision has changed it.

3. Check the level: IR, AMC or GM

An implementing rule is the requirement, AMC is an acceptable means of meeting it, and GM explains. The guide to reading EASA rules sets out how the layers relate. An answer that blends them turns "should" into "shall".

In the example, the implementing rule gives no interval at all. How the interval is set sits one level down, in the AMC:

AMC 145.A.40(b), point 2 · Easy Access Rules for Continuing Airworthiness, 2025-09-02 revision

  1. Inspection, service or calibration on a regular basis should be in accordance with the equipment manufacturers' instructions except where the organisation can show by results that a different time period is appropriate in a particular case.

Neither level says 12 months.

4. Find the conditions and exceptions

Look for "unless", "except", "where", "as appropriate" and defined terms. They decide who a rule applies to and when. An answer that drops an exception reads stricter than the rule. One that drops a condition reads looser. Both put the wrong test into a finding.

The example has three: "as appropriate" in the rule, the "except where the organisation can show by results" exception in the AMC, and a definition in point 3 of the same AMC:

  1. In this context officially recognised standard means those standards established or published by an official body whether having legal personality or not, which are widely recognised by the air transport sector as constituting good practice.

"A national standard" is not that test.

5. Check the sentence says what the passage says

Put them side by side and go phrase by phrase. Every number, duration, "must" and "all" in the answer should appear in the passage. If it does not, it came from somewhere else.

The answer saysThe passage says
"all tools and test equipment""all tools, equipment and particularly test equipment, as appropriate"
"at least every 12 months""at a frequency to ensure serviceability and accuracy" (no interval)
"a national standard""an officially recognised standard", defined in AMC 145.A.40(b) point 3
"calibration records must be kept""Records of such calibrations and traceability to the standard used shall be kept"

Every clause drifts, and one adds a number the rule does not contain. The finding at ExampleMRO cites 145.A.40(b) as written and tests the wrench against the interval in ExampleMRO's own tool control procedure. Setting those intervals and writing the finding are covered in tool control and calibration under 145.A.40.

The rule to keep from this page: a citation that exists can still fail to support the claim. It is the misgrounded answer in the Stanford study, and step 5 is how you catch it.

This is the check for one cited sentence. The wider routine, from your exposition to who reviews the draft, is in Can you use ChatGPT for EASA compliance work?. Which documents carry weight is in why aviation AI must cite approved sources.

Doing this in Avioverse

Start with the passages, not an answer. Switch the message box from Ask Metis to Find sources. It returns source excerpts with no generated answer and uses no AI credits. A new message box searches with Words & meaning; on the Free plan it falls back to Words & references and says so. Type a rule number such as 145.A.40 and it looks that reference up directly. Cards show the version details the library holds: amendment, status, issue date, stored applicable-from and applicable-to dates, and whether it is the current or a historical stored version. Filter by Requirements (IR), Acceptable means (AMC) or Guidance (GM), and open Related clauses for cited provisions, related AMC/GM and what cites the point. Compare selected puts two or more passages side by side. Copy brief copies the passages, their version details and your note; Save to Brain keeps the brief on a paid plan. Ask Metis about selected pre-fills a question and uses your allowance only when you send it.

When you ask for an answer, select a citation number to see an extract of its source. Open takes you to that exact provision in the regulation reader. Source check re-reads a reply that used sources against the evidence retrieved for it. It runs on every plan, including Free. On a paid plan you can turn it off under Intelligence (the button shows your current mode, such as Standard) → Answer checking, and your last choice becomes the default for new chats. A checked reply carries a label:

  • Checked against sources: the check finished with no unresolved issue against the retrieved evidence. It does not prove every relevant source was found.
  • Checked against context: the same, against the available context (such as what you told Metis) when the reply shows no numbered sources.
  • Review suggested: the check flagged a possible problem that is still unresolved.
  • Not checked: the full check did not finish, or Source check was off.

Treat a reply with no label as unchecked. A label is an automated reading of the same material, not an approval: the five steps are still yours.

Frequently asked questions

Can you trust ChatGPT on EASA regulations?

Use its answer as a pointer to where the rule is, not as the rule. Studies of cited AI answers in law, medicine and news keep finding citations that do not support the sentence they sit on. Open the cited passage and compare it with the answer before you use it.

What is an AI hallucination in regulatory work?

An answer that is wrong, or that is not supported by the source it cites. The Stanford RegLab study of legal AI tools counted a real citation that does not support the claim as a hallucination, which is the failure that matters most in a finding.

How do I check an AI answer against the regulation?

Open the cited passage itself, confirm it is the revision in force, check whether it is IR, AMC or GM, find the conditions and exceptions, and compare the answer with the passage word by word. Every number and "must" in the answer should appear in the passage.

If the AI gives a rule number, is the answer right?

Not necessarily. A citation that exists can still fail to support the claim. An answer can name the right paragraph and add an interval, a standard or a scope that the paragraph does not contain.

Does EASA regulate AI assistants such as ChatGPT?

We found no adopted EASA rule for AI assistants as of 25 September 2026. NPA 2025-07 is a proposal that does not yet fully cover generative AI, and the AI concept paper is guidance. Neither makes an AI answer a source you can cite.

Does Checked against sources in Avioverse mean the answer is correct?

No. It means the automated check finished with no unresolved issue against the evidence that reply retrieved. It does not prove that every relevant source was found, so you still open the passage.

Related

Written by Dionysis Kefalas. Retired Hellenic Air Force Captain and founder of Avioverse. About the author

Request early access →

Metis prepares answers from the EASA regulation library with numbered sources you can open, so you check the rule text before you rely on it. Opens in October 2026.

ShareLinkedInX