In the library of a Swiss university department sits a pile of rejected bachelor’s theses from the past twelve months. Three of them share the same punchline: a key citation points to an article that doesn’t exist. The author and title sound credible, the page number is precise — and yet the source does not exist. The students asked ChatGPT, ChatGPT answered plausibly, and the entry went straight into the thesis. The examiners tried to look up the article. That’s when it came to light.
This post is a collection of seven concrete detection methods for spotting hallucinated sources in your own thesis before you submit it. No theory essay — seven tests, each taking 30 seconds to five minutes, illustrated with real patterns from student work (anonymised).
If you don’t yet know why language models fabricate sources, there’s a dedicated technical post for that: AI source hallucinations — how and why they happen. This post is about what you as a student can do about it.
What hallucinated sources are — briefly
A hallucinated source is a bibliographic reference that a language model has fabricated. Three variants turn up in student work:
- Fully fabricated. Neither the author, title, nor publisher exists — or they don’t exist in this combination.
- Partially fabricated. Author and publisher are real, but the title or edition doesn’t exist; or the page number is off by 50 pages.
- Misattributed. The source exists but doesn’t say what is cited — sometimes the opposite.
Tests 1–6 catch the first two types. The third requires the content check in Test 7. Anyone who has used ChatGPT as a research assistant should run all three, to be safe.
1. The DOI test (10 seconds per citation)
Every serious academic publication today has a Digital Object Identifier (DOI). In ChatGPT-generated source suggestions, the DOI is either absent entirely — or fabricated.
How to test:
- Take the DOI from the suggested reference (format:
10.xxxx/yyyy). - Append it to
https://doi.org/, e.g.https://doi.org/10.1234/zbp.2019.47.3.234. - Press Enter. If it resolves to the stated paper, the source is real. If it returns a 404, leads to a different paper, or lands on a generic publisher page, the source is hallucinated.
Real pattern from a bachelor’s thesis (anonymised):
Mueller, J. (2019). Social cognition and trust formation in digital learning groups. Journal of Educational Psychology, 47(3), 234–251. DOI: 10.1024/1010-0652/a000178
The DOI looks plausible — Hogrefe format, structurally correct. Does it resolve? Type it in. Three times out of four it leads nowhere or to a completely different paper from 2014. If the DOI is live, also check the volume and year: Hogrefe’s educational psychology journal was only at volume 33 in 2019 — volume 47 never existed.
Watch out: Some genuine sources from the 1990s genuinely have no DOI. A missing DOI is therefore not automatically suspicious. However: a missing DOI on a paper from 2010 onwards is unusual — and ChatGPT often leaves the field empty in hallucinations to avoid detection. More on what belongs in a reference: DOI, ISBN, ISSN — what a citation needs.
2. Google Scholar test with quotation marks (30 seconds)
For every ChatGPT-suggested source, take the exact title in quotation marks and search for it on Google Scholar (scholar.google.com).
What you see:
- Multiple hits with the same title → source exists, investigate further.
- Exactly one hit with a different title that shares only a few words → hallucinated.
- Zero hits in Google Scholar AND zero hits in a regular Google search → very likely fabricated.
Academic literature leaves traces: citations, publisher pages, library catalogues, Google Books snippets. A book that has supposedly existed since 2017 but whose title appears nowhere online in 2026 does not exist.
Real pattern:
A student receives the source “Schmidt, K. (2021). Foundations of Quantitative Social Research in Educational Economics. Springer VS.” from ChatGPT. Plausible, right? Schmidt is a common surname, the topic sounds clean, and Springer VS publishes exactly these kinds of books. In Google Scholar with quotation marks: zero hits. In the DNB catalogue: no entry. In the Springer catalogue: no entry. Fabricated.
Watch out: Very recent sources from 2025 or 2026 are sometimes not yet fully indexed. If the title is new and Scholar finds nothing, also check the publisher directly. If the publisher’s catalogue has no entry either, the source is gone.
3. Your university library catalogue (1 minute)
For books (monographs, edited volumes), the fastest test is your university library catalogue. Swiss universities use SLSP / Swisscovery, German ones typically KIT Karlsruhe or local catalogues, Austrian ones the OBV. Search by author + keyword.
What you expect: a real book → a hit with call number, location, and availability. Sometimes even a direct link to the e-book.
What you don’t expect: “No results found” for a supposed standard text from a German-language academic publisher from 2018. Mohr Siebeck, Suhrkamp, De Gruyter, Vandenhoeck & Ruprecht, Springer VS — all these publishers are catalogued at every university library. A gap means: it doesn’t exist.
Real pattern:
Hofmann, R. (2020). Discourse Analysis in German Political Science. Suhrkamp, p. 187.
Rebecca Hofmann exists, is a political scientist, but has never published with Suhrkamp (Suhrkamp publishes fiction and philosophy, not political science textbooks). Publisher confusion — a classic ChatGPT error. No hit in Swisscovery. No hit in the DNB catalogue. Hallucinated.
4. ISBN checksum and DNB entry (1 minute)
The ISBN-13 is mathematically verifiable. The last digit is a check digit. A fabricated ISBN often fails the algorithm already.
Quick online test: https://www.isbn-check.de/ or any ISBN validator. Paste it in and check.
In addition: Search the ISBN at the Deutsche Nationalbibliothek. Any German-language book must appear there — legal deposit. No DNB entry means no book.
Real pattern:
ISBN 978-3-540-12345-6
Check digit: if you run the algorithm (alternating multiplication of digits by 1 and 3, modulo 10), the check digit 6 does not come out. Fabricated ISBN. Even if the check digit happens to be correct — the DNB has no entry. Dead.
For what an ISBN actually encodes and which format applies when, see DOI, ISBN, ISSN — what a citation needs.
5. Page number sanity check (30 seconds)
ChatGPT fabricates page numbers without knowing the actual length of the book. Classic error: citing “p. 234” from a work that has 180 pages in total.
How to test:
- Search for the book in your library catalogue or on Google Books — most hits show the total page count.
- Compare this to your cited page number.
- If your page falls past the end of the book or in the index, the citation is hallucinated or at least placed incorrectly.
Real pattern:
ChatGPT cites “Bourdieu, P. (1982). Distinction. Suhrkamp, p. 612.” The book exists (good), and in the Suhrkamp paperback edition it runs to 957 pages (so p. 612 would be plausible), but the first hardback edition only runs to about 800. However: when you open that page, it covers a different topic from what the student claims. That’s variant 2 — the source exists, but the page is wrong. What you need to read against the original is covered in the citation-checking checklist for bachelor’s theses.
Watch out: Different editions have different page numbers. If you cite the second edition from 2007 but someone checks the first edition from 1982, the pages won’t match. Always include the exact edition in your reference.
6. Journal volume and year cross-check (1 minute)
For journal articles, there is a dry but reliable check: does the stated volume and issue number match the year of publication?
Every established journal publishes one volume per year (some per half-year). From a volume number and year you can reconstruct whether the issue actually exists. Example: the Journal of Sociology was at volume 53 in 2024, 52 in 2023, 51 in 2022, and so on. If ChatGPT gives you “Vol. 67, Issue 4 (2024)”, that’s fabricated — volume 67 wouldn’t be reached until 2038.
How to test:
- Go to the publisher’s page for the journal (or JSTOR, ScienceDirect, SAGE Journals).
- Look up the volume for the stated year.
- Check whether the issue number and article title match.
Real pattern:
Weber, S., & Lang, T. (2018). Resilience as an education policy concept. Journal of Pedagogy, 64(6), 821–839.
The Journal of Pedagogy exists, in 2018 it really was at vol. 64, it publishes 6 issues per year — everything structurally correct. In the Beltz archive for 2018, issue 6: different authors, different topics, this article is not there. Hallucinated. ChatGPT filled in the right schema but invented the content.
Special case — predatory journals: Sometimes the journal does exist, but it’s predatory — scientifically worthless. That’s a different class of problem, covered in detail in Identifying predatory journals.
7. The substantive original-text check (2–5 minutes per citation)
The first six tests expose existence hallucinations — sources that don’t exist. The most treacherous variant, however, is the content hallucination: the source exists, but it doesn’t say what you claim.
How to test:
- Get hold of the source PDF (university library, interlibrary loan, contact the author — guidance in Checking citations in your bachelor’s thesis).
- Go to the cited page. Read the surrounding passage.
- Compare word for word for direct quotations, meaning for meaning for paraphrases.
Real pattern:
ChatGPT suggestion: “According to Habermas (1981), communicative action is a purely instrumental form of social interaction.” Habermas exists, the book (The Theory of Communicative Action) exists. But: Habermas argues in 1981 the exact opposite — communicative action stands in explicit opposition to instrumental rationality. ChatGPT correctly named the standard work and inverted the argument. Classic content hallucination.
What to remember: For every key citation in your thesis, this is the test that cannot be skipped. Tests 1–6 check the form; Test 7 checks the substance. If you’re only going to check 30 citations for substance, check the 30 most important ones — the ones your argument rests on.
A note on secondary citations: don’t just check what the intermediate author says — look at the original. Sometimes the intermediate author already cited it incorrectly. How to handle this cleanly is covered in Using secondary citations correctly.
What to do if you find a hallucination
Three options, in this order:
- Find a real source. If the cited claim is substantively correct, some real author has certainly made it before. Find the real source and replace the hallucinated one.
- Rephrase the claim so it needs no citation. If the point is common knowledge (e.g. “language models generate one token at a time”), you can leave it unsupported.
- Delete the claim. If you find no source and the claim is central to your argument, it was speculative to begin with. Cut it.
What you don’t do: leave the hallucinated source in and hope nobody checks. Examiners today routinely pick three to five random citations from every thesis and look them up. A fabricated source leads, in the mild case, to a requirement to resubmit; in serious cases it is treated as an attempt to deceive. The matter then goes to the examination board — which can sanction up to revocation of the degree. More on the legal side: Declaration of academic integrity — what happens if you violate it.
How to prevent hallucinations from the start
The most efficient prevention is a clean separation: ChatGPT for language, libraries and databases for sources.
- Brainstorming, outlining, language editing: yes, with ChatGPT.
- Searching for academic literature, getting source suggestions, building a bibliography: not with ChatGPT — use Google Scholar, your library catalogue, Web of Science, JSTOR.
- If you do use ChatGPT as a starting point for research: run every suggestion through Tests 1–6 before importing it into Zotero.
For a summary of what AI use is generally permitted in academic theses in 2026, see Citing ChatGPT as a source — what universities allow in 2026.
When does an automated tool make sense?
Manual checking is feasible but tedious. For 80 citations in a bachelor’s thesis, Tests 1–6 take about two hours; Test 7 (content) adds another three to six. Anyone who stumbles across this in the wrong moment of the final stretch knows the feeling: one hallucination is enough to delay submission.
That’s exactly what Acurio is built for: you upload your DOCX and your source PDFs, and multiple language models check every claim against the actual passage in the original text. For each citation you get back: supported, partially supported, unsupported — with a source excerpt and confidence score. This covers Tests 5–7; Tests 1–4 (existence, DOI, ISBN, volume) remain your job, but with Zotero they take minutes.
For a direct comparison of different tools, see Checking your thesis — a tool comparison.
Summary
Hallucinated sources are the most common new problem in student theses in 2026 — and the one most reliably triggering an examination procedure if discovered. The seven tests in this post take a total of a few minutes per citation. The investment is always worth it; the only case where skipping them is not costly is when the examiners also don’t check — and you shouldn’t count on that.
The rule of thumb: any source that doesn’t come from a database, a library catalogue, or a journal with a working DOI gets checked against at least two of Tests 1–6. If both come back negative, the source goes.
Anyone who has once been caught with a hallucinated source in their thesis checks everything for the rest of their degree. It’s easier to do that from the start. Acurio cuts the checking time to under an hour — leaving the rest of your final stretch in peace.