A New York lawyer asked ChatGPT whether the court cases it gave him were real. It said yes. They weren't. Why chatbots invent facts in a confident voice, what OpenAI's own researchers argue, and the one check that would have saved him.
Takeaways
- A chatbot is built to predict the next word, and, OpenAI's own researchers argued in 2025, the way we test these machines rewards a confident guess over "I don't know".
- The rarer and more specific the fact, the more the answer is a guess: in one 2023 test, 55% of references from GPT-3.5 and 18% from GPT-4 were invented.
- This week: sort one question into pattern or fact, and give the fact one check from outside the machine.
Chapters
- Six cases that never existed0:17
- A super search engine?2:03
- Graded like a test-taker4:04
- Same question, different stakes7:09
- Hear it done8:39
- Build yours9:57
- The answer11:37
References
- United States District Court, S.D.N.Y. (Castel, J.) — Mata v. Avianca, Inc., No. 22-cv-1461 (PKC), Opinion and Order on Sanctions(2023)
- ABA Journal — Lawyers who 'doubled down' and defended ChatGPT's fake cases must pay $5K, judge says(2023)
- Courthouse News Service — Lawyer who cited bogus legal opinions from ChatGPT pleads AI ignorance(2023)
- Shannon, C. E. — Prediction and entropy of printed English(1951)
- OpenAI — GPT-4 Technical Report(2023)
- Kalai, A. T., Nachum, O., Vempala, S. S., & Zhang, E. — Why language models hallucinate(2025)
- Walters, W. H., & Wilder, E. I. — Fabrication and errors in the bibliographic citations generated by ChatGPT(2023)
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. — Hallucination-free? Assessing the reliability of leading AI legal research tools(2025)
- Charlotin, D. — AI Hallucination Cases Database(2026)