← Back to Research Radar
Academic Publication Academic Publication

Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis

291
Citations
May 22, 2024
Published Date

Research Abstract & Technology Focus

Background
Large language models (LLMs) have raised both interest and concern in the academic community. They offer the potential for automating literature search and synthesis for systematic reviews but raise concerns regarding their reliability, as the tendency to generate unsupported (hallucinated) content persist.


Objective
The aim of the study is to assess the performance of LLMs such as ChatGPT and Bard (subsequently rebranded Gemini) to produce references in the context of scientific writing.


Methods
The performance of ChatGPT and Bard in replicating the results of human-conducted systematic reviews was assessed. Using systematic reviews pertaining to shoulder rotator cuff pathology, these LLMs were tested by providing the same inclusion criteria and comparing the results with original systematic review references, serving as gold standards. The study used 3 key performance metrics: recall, precision, and F1-score, alongside the hallucination rate. Papers were considered “hallucinated” if any 2 of the following information were wrong: title, first author, or year of publication.


Results
In total, 11 systematic reviews across 4 fields yielded 33 prompts to LLMs (3 LLMs×11 reviews), with 471 references analyzed. Precision rates for GPT-3.5, GPT-4, and Bard were 9.4% (13/139), 13.4% (16/119), and 0% (0/104) respectively (P
Read Full Literature

AI Semantic Synergy Context

Connecting this academic literature to real-world market discussions and products.

news.ycombinator.com › AI insight
0%

Show HN: Large scale hallucinated citation problem in published literature

Grounded AI's study exposes a critical integrity crisis in academic publishing, directly linked to generative AI. The estimated 'hundreds of thousands of papers affected in 2025' highlights a syste...

roipad.com › trend story
0%

New Study Raises Concerns About AI Chatbots Fueling Delusional Thinking

"Emerging evidence indicates that agential AI might validate or amplify delusional or grandiose content, particularly in users already vulnerable to psychosis," writes Dr Hamilton Morrin, a psychia...

crossref.org › academic paper
0%

Detecting hallucinations in large language models using semantic entropy

AbstractLarge language model (LLM) systems, such as ChatGPT1or Gemini2, can show impressive reasoning and question-answering capabilities but often ‘hallucinate’ false outputs and unsubstantiated a...

news.ycombinator.com › discussion
0%

Show HN: Large scale hallucinated citation problem in published literature

Hey, Nick Morley from Grounded AI here (https://groundedai.company/)We collaborated with Nature to study the extent of fake/frankenstein citations in scholarly literature (from top 5 publishers - S...

roipad.com › trend story
0%

ChatGPT Health Underestimates Medical Emergencies, Study Finds

It is also inconsistent with suicide-risk alerts, the researchers said.

Frequently Asked Questions (FAQ)

Curated market intelligence mapped to this research.

What is the core focus of the research titled 'Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis'?

This literature focuses on: Background Large language models (LLMs) have raised both interest and concern in the academic community. They offer the potential for automating literature search and synthesis for systematic reviews but raise concerns regardin...

Are there open-source GitHub repositories related to Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis?

Yes, open-source projects like anthropics/claude-desktop-buddy (Reference and an example for the Bluetooth API for makers in Claude Cowork & Claude Code Desktop) are actively building upon these concepts.

How is the concept of 'Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis' being discussed by engineers on Hacker News?

Yes, highly correlated activity was mapped. An entry titled 'Show HN: Large scale hallucinated citation problem in published literature' discusses this: Grounded AI's study exposes a critical integrity crisis in academic publishing, directly linked to generative AI. The estimated 'hundreds of thousa...

Are there commercial applications of 'Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis' in market news publications?

Yes, highly correlated activity was mapped. An entry titled 'New Study Raises Concerns About AI Chatbots Fueling Delusional Thinking' discusses this: "Emerging evidence indicates that agential AI might validate or amplify delusional or grandiose content, particularly in users already vulnerable t...

What other academic literature is closely related to 'Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis'?

Yes, highly correlated activity was mapped. An entry titled 'Detecting hallucinations in large language models using semantic entropy' discusses this: AbstractLarge language model (LLM) systems, such as ChatGPT1or Gemini2, can show impressive reasoning and question-answering capabilities but often...

Cite this Market Intelligence Report

Reference our AI-mapped synergy between this research and the commercial market to instantly build authority.

Commercial Realization

Startups and Open Source tools heavily associated with the concepts explored in this paper.

Associated Media Narrative