Academic Publication Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis
Research Abstract & Technology Focus
Large language models (LLMs) have raised both interest and concern in the academic community. They offer the potential for automating literature search and synthesis for systematic reviews but raise concerns regarding their reliability, as the tendency to generate unsupported (hallucinated) content persist.
Objective
The aim of the study is to assess the performance of LLMs such as ChatGPT and Bard (subsequently rebranded Gemini) to produce references in the context of scientific writing.
Methods
The performance of ChatGPT and Bard in replicating the results of human-conducted systematic reviews was assessed. Using systematic reviews pertaining to shoulder rotator cuff pathology, these LLMs were tested by providing the same inclusion criteria and comparing the results with original systematic review references, serving as gold standards. The study used 3 key performance metrics: recall, precision, and F1-score, alongside the hallucination rate. Papers were considered “hallucinated” if any 2 of the following information were wrong: title, first author, or year of publication.
Results
In total, 11 systematic reviews across 4 fields yielded 33 prompts to LLMs (3 LLMs×11 reviews), with 471 references analyzed. Precision rates for GPT-3.5, GPT-4, and Bard were 9.4% (13/139), 13.4% (16/119), and 0% (0/104) respectively (P
AI Semantic Synergy Context
Connecting this academic literature to real-world market discussions and products.
Show HN: Large scale hallucinated citation problem in published literature
Grounded AI's study exposes a critical integrity crisis in academic publishing, directly linked to generative AI. The estimated 'hundreds of thousands of papers affected in 2025' highlights a syste...
New Study Raises Concerns About AI Chatbots Fueling Delusional Thinking
"Emerging evidence indicates that agential AI might validate or amplify delusional or grandiose content, particularly in users already vulnerable to psychosis," writes Dr Hamilton Morrin, a psychia...
Detecting hallucinations in large language models using semantic entropy
AbstractLarge language model (LLM) systems, such as ChatGPT1or Gemini2, can show impressive reasoning and question-answering capabilities but often ‘hallucinate’ false outputs and unsubstantiated a...
Show HN: Large scale hallucinated citation problem in published literature
Hey, Nick Morley from Grounded AI here (https://groundedai.company/)We collaborated with Nature to study the extent of fake/frankenstein citations in scholarly literature (from top 5 publishers - S...
ChatGPT Health Underestimates Medical Emergencies, Study Finds
It is also inconsistent with suicide-risk alerts, the researchers said.
Frequently Asked Questions (FAQ)
Curated market intelligence mapped to this research.
What is the core focus of the research titled 'Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis'?
This literature focuses on: Background Large language models (LLMs) have raised both interest and concern in the academic community. They offer the potential for automating literature search and synthesis for systematic reviews but raise concerns regardin...
Are there open-source GitHub repositories related to Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis?
Yes, open-source projects like anthropics/claude-desktop-buddy (Reference and an example for the Bluetooth API for makers in Claude Cowork & Claude Code Desktop) are actively building upon these concepts.
How is the concept of 'Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis' being discussed by engineers on Hacker News?
Yes, highly correlated activity was mapped. An entry titled 'Show HN: Large scale hallucinated citation problem in published literature' discusses this: Grounded AI's study exposes a critical integrity crisis in academic publishing, directly linked to generative AI. The estimated 'hundreds of thousa...
Are there commercial applications of 'Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis' in market news publications?
Yes, highly correlated activity was mapped. An entry titled 'New Study Raises Concerns About AI Chatbots Fueling Delusional Thinking' discusses this: "Emerging evidence indicates that agential AI might validate or amplify delusional or grandiose content, particularly in users already vulnerable t...
What other academic literature is closely related to 'Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis'?
Yes, highly correlated activity was mapped. An entry titled 'Detecting hallucinations in large language models using semantic entropy' discusses this: AbstractLarge language model (LLM) systems, such as ChatGPT1or Gemini2, can show impressive reasoning and question-answering capabilities but often...
Cite this Market Intelligence Report
Reference our AI-mapped synergy between this research and the commercial market to instantly build authority.
Commercial Realization
Startups and Open Source tools heavily associated with the concepts explored in this paper.
-
GitHubanthropics/claude-desktop-buddy
-
GitHubmark9-droid/TomodachiPC
Associated Media Narrative
- Osaurus, the Native macOS Harness for Local AI Models, Celebrates 7.3K GitHub Stars and #2 Product of the Day on Product Hunt
- GPT-Red beat human red teamers on a prompt injection test
- A pilot study on the impact mechanism of internal and external leading variables on consumers’ purchase intention and healthy dietary behavior of plant-rich foods
SaaS Metrics