“78 percent of students use ChatGPT.” Nice sentence, fits neatly into the introduction, comes with a footnote. Your examiner at the viva asks: “78 percent of whom, measured when, and what counts as use?” If you can’t answer, the number collapses — and half your argument goes with it. This guide walks through how to cite a statistic so it survives the follow-up questions: from tracking down the primary source and avoiding the cascade trap to what has to sit next to the number.
A statistic is a claim, not proof
A number in your thesis is nothing more than a sentence with numerical precision. “78 percent” makes the same kind of claim as “most people” — just with more perceived authority. That authority is earned, through a source that actually backs the number.
Concretely, that means every statistic in your running text needs a reference that answers three questions, or it isn’t formally verifiable:
- Who measured this (institution, author, publication year)?
- What exactly was measured (definition of the metric, reference period, target population)?
- How large is the sample (n=), and how representative is it?
If one of those three is missing, the number is at best a hunch with a decimal point.
The cascade trap — how numbers rot in the citation chain
The most common mistake in theses isn’t the wrong number, it’s the wrong source. Here’s how a statistic typically travels:
- The study: a research institute runs a survey of 812 students between January and March.
- The press release: their comms team writes “Nearly 80 percent of students use AI tools during their studies.” (No n, no time window, no definition of “use”.)
- The news article: a newspaper picks up the number and credits the institute but doesn’t link the study.
- The blog: an SEO blog paraphrases the news article and rounds to “80 percent”.
- Your thesis: you cite the blog — and, implicitly, the three-times-reshaped number.
Every hop loses context: sampling window, definition, sample size, confidence interval. What you end up citing is a claim nobody in the original actually made.
Rule of thumb: go as far upstream toward the primary source as you can. A McKinsey report is not a primary source when it leans on an OECD analysis that is itself a meta-analysis. Click through until you reach the study that collected the number, not the one that repeated it.
How to actually find the primary source
Four moves that work in nine cases out of ten:
- Read the footnote of the secondary source. Actually read it. Reports and papers usually put the references at the end, not in the paragraph.
- Google Scholar with author + fieldwork year. “Meier 2022 digital students survey” filters out the secondary citations.
- Follow the DOI. If the secondary citation gives a DOI, it takes you to the primary source. For whitepapers, often straight to the PDF.
- Scan the institution’s website. Big data collectors (Pew Research, Ofcom, Eurofound, Bertelsmann Stiftung, gfs.bern) keep study archives. Search their site by title and year — it turns up.
If you truly can’t find the primary source, cite it cleanly as a secondary citation: “Meier & Schulz (2022), as cited in Federal Education Report (2024, p. 45).” That’s more honest than pretending you read the original.
What has to sit next to the number
A good statistical reference answers more than “who said it”. Six fields make the difference:
| Field | Example |
|---|---|
| Reference period | “January – March 2023” (not “2024” for the publication) |
| Target population | “Students at German universities” |
| Sample size | n = 812 |
| Collection method | Online panel, stratified random sample |
| Metric definition | “Use at least once per week” |
| Margin of error / confidence interval | ±3.4 percentage points at 95% confidence |
You don’t have to squeeze all of it into the citation — that would be unreadable. But the pieces load-bearing for your argument belong in the running text or a footnote:
An online survey of 812 students at German universities (fieldwork Q1 2023, ±3.4 pp) found that 78% use ChatGPT at least once a week for study-related work (Meier & Schulz, 2023).
That sentence holds up under follow-up questions. “78% (Meier & Schulz, 2023)” holds up under nothing.
Percentage points, percent, growth rates — the classic mistakes
A large share of misunderstandings isn’t about wrong sources, it’s about wrong arithmetic on correct numbers. The three most common traps:
1. Percent vs. percentage points. When a share rises from 5% to 10%, that’s +5 percentage points or +100% relative increase — depending on how you phrase it. “The share went up by 5 percent” is wrong if you actually mean +5 pp. Rule of thumb: absolute change between two shares → percentage points, relative change → percent.
2. Comparing across incompatible periods. “Revenue is up 15% versus 2019” quietly hides that 2020 and 2021 may have crashed. If your baseline year had a special effect (pandemic, crisis, regulatory change), name it or pick a different base year.
3. Doubled, halved, quadrupled. Sounds dramatic but is often misleading at low base values. “Cases doubled” from 2 to 4 per million is statistically weak. Always give the absolute values.
When Statista, McKinsey or BCG is your only lead
Statista is not a data source, it’s an aggregator — every Statista chart names its primary source at the bottom. Cite that primary source, not Statista. If the primary source is behind a paywall, Statista is the trail back to your university library: grab author, title, year and pull it via the university catalogue.
Consulting reports (McKinsey Global Institute, BCG, Deloitte) are stronger than their reputation when they contain original fieldwork — then they are primary sources. When they synthesise other people’s numbers, they’re secondary sources with editorial polish. Check the methods appendix: “own analysis of dataset X”? Primary. “Based on public sources”? Secondary.
For the reference itself, the usual report conventions apply. In APA 7:
McKinsey Global Institute. (2024). The economic potential of generative AI: The next productivity frontier. McKinsey & Company. https://www.mckinsey.com/mgi/our-research/…
Two details: corporate authors go into the bibliography without a first name, and the institute’s full name is spelled out — “MGI” alone doesn’t cut it there.
When you do the maths yourself — “own calculation”
The moment you derive a new number from a source (mean, growth rate, share, ratio), it isn’t a secondary statistic anymore, it’s your analysis. Flag it:
The five-year average sits at 12.3% (own calculation based on Eurostat, table demo_pjan).
Two rules:
- Only as many decimals as the input justifies. If your inputs are rounded to whole percentages, “12.3456%” is false precision. Round to the coarsest precision of your raw data.
- Keep the calculation reproducible. An Excel sheet or R script in the appendix or a repository — not a citation obligation, but a reproducibility obligation.
The three-question test for every statistic
Before you commit a number to running text, run it through:
- Is it really phrased that way in the original? Pull up the title, table, page number and compare. In reports it’s often the executive summary vs. the detail chapter that diverge.
- Does it match your subject matter in time? A number from 2018 doesn’t prove a claim about 2024, even if it’s the only one available.
- Do the source and your sentence use the same definition? “Employed” under ILO is not the same as “employees subject to social security contributions” under a national employment agency. If your text uses one term, the source has to measure it too.
Anyone who applies this test to every cited statistic will delete painful amounts of numbers while writing — and dodge the awkward follow-ups at the viva.
Where Acurio comes in
The formal side of citing a statistic is easy: author, year, title, locus. The second layer is content: does the source actually say what you claim? Did you take the right number from the right table for the right period? These are the errors that get caught in the defence.
For text citations, Acurio handles that content check: you upload your thesis together with the primary sources, and every cited number is compared with the original. For datasets, the methods section stays your tool — a reproducible script that derives every headline figure from the raw data. Combine the two and you walk into the viva with a clear head.
A statistic that holds up in a thesis rests on three legs: an identifiable primary source, enough context in the text (n, period, definition), and a clean calculation when you derive numbers yourself. Kick one of those legs out and the number falls — and takes the passage that depended on it with it.