Three years ago, AI-assisted literature research meant asking ChatGPT for studies and hoping they existed. Mostly, they didn’t. Since then, a whole market segment has taken shape — specialised tools that sit on top of a scientific full-text corpus and return answers with citations. Four have broken out: Elicit, SciSpace, Consensus, and Research Rabbit. This post compares them by what actually helps you during a bachelor or master thesis — and where each one gets it wrong.
What these tools do differently from ChatGPT
The decisive difference is the data behind them. A general language model like ChatGPT has seen papers, but it can neither ground itself in a specific full-text nor verify a DOI. You know the result: hallucinated sources that look impeccable and don’t exist in any catalogue.
AI research tools sit on top of a closed corpus — usually Semantic Scholar (around 220 million publications), CORE, OpenAlex, or a proprietary aggregation. Your question is semantically matched against that pool, and the answers come with title, authors, year, and a link to the original. Fabricated studies are essentially off the table. What’s left as an error source: the tool’s summary of the paper. And that can be off.
If entries from AI research have already made it into your bibliography, a counter-check before submission is worth it: the free Quick Check matches every entry against catalogues and databases, no sign-up needed. How the check works in detail is on the bibliography checker page.
Elicit — the extraction engine
Elicit came out of the AI research spin-off Ought and positions itself as a “research assistant”. The core is a table: you enter a research question, the tool searches Semantic Scholar, and it opens a row for each paper found — with title, abstract, and, more importantly, columns you define yourself.
Example: you ask “How does remote work affect knowledge-worker productivity?” Elicit searches, shows 20 relevant papers, and you add columns like “sample size”, “study design”, “effect direction”, or “country”. For each row, the tool extracts the answer straight from the full text, with a quote and page reference.
Strength: no other tool brings structure to a messy research landscape this fast. If you’re building a systematic overview or a state-of-the-art section, you save days.
Weakness: the extraction is only as good as the model behind it. On complex methodological questions — “which control variables were used?” — Elicit regularly gets it wrong because it takes the question too literally. You verify, every time.
Cost: free tier with limited searches per month, paid from around USD 12/month.
SciSpace — chat with the paper
SciSpace (formerly Typeset) takes a different route. Instead of tables, it offers a chat where you talk to one or more papers. You upload a PDF or pick from the corpus, and you ask questions: “What’s the main claim?”, “Which method did they use?”, “How does this differ from study X?”
SciSpace also has general search, a paraphrasing mode, and a citation generator — the three extras are peripheral. The chat is the core idea.
Strength: when you need to understand a study fast, SciSpace is unbeatable. The chat finds hidden details in long papers you’d have skipped on a first pass — the exact formula for the significance test used, the definition of a poorly named construct, the location of data collection.
Weakness: the chat tempts you not to read the paper. That is exactly the mistake. SciSpace tells you what’s in it — but whether you can position it correctly without reading it yourself is a different question. In an exam situation, worthless.
Cost: free version to get started, premium from around USD 12/month.
Consensus — the opinion detector
Consensus asks a different kind of question. Not “give me papers on this topic”, but “what does the research say about this specific yes-or-no question”. You type, for example, “Does remote work increase productivity?” Consensus searches the corpus and returns a list of matching studies with short excerpts. Prominently at the top: a “Consensus Meter” — bars showing what percentage of studies say yes, no, or mixed.
Strength: the Consensus Meter is a good reality check on claims that circulate as “well established”. If you assume remote work clearly raises or lowers productivity, a 60/30/10 split will straighten you out fast.
Weakness: the yes/no classification is coarse. A study finding a small effect in a niche population gets the same weight as a large meta-analysis. Useful as a starting point, too blunt for a state-of-the-art review. And: without the full text behind the excerpt, you don’t know whether the study even applies to your context.
Cost: free tier for a limited number of questions per month, premium from around USD 9/month.
Research Rabbit — the map
Research Rabbit has no chat, no table, no meter. It shows you relationships. You throw in one or more papers you already know, and Research Rabbit builds a graph out of them: which papers cite this one, which are cited by it, which come from the same authors, which were published around the same time.
Strength: for discovering literature you’d never have found otherwise, the tool is excellent. If you’re starting in a new field and you have a “seed” paper, Research Rabbit makes the relevant neighbourhood visible — including the authors who keep coming up.
Weakness: Research Rabbit gives you no substantive analysis. It says “these papers are related” — what they say is on you. And the graph gets crowded fast if you cast a wide net.
Cost: free, no subscription.
Direct comparison
| Tool | Core idea | Strongest at | Weakness | Cost |
|---|---|---|---|---|
| Elicit | Extraction table | Systematic overview, state of the art | Extraction shaky on complex methods | Freemium, from ~USD 12 |
| SciSpace | Chat with the paper | Understanding a single study fast | Tempts you to skip the read | Freemium, from ~USD 12 |
| Consensus | Yes/no consensus | Reality check on everyday claims | Coarse classification | Freemium, from ~USD 9 |
| Research Rabbit | Citation graph | Discovering related literature | No substantive analysis | free |
These tools don’t exclude each other. In practice you combine two or three — each for the phase where it’s strongest.
A realistic workflow
Take a bachelor thesis on “remote work and productivity”. Here’s how a research workflow might use the four tools productively:
- Consensus — rough overview of the study landscape. Quickly shows the research is mixed and that different domains (knowledge work vs. manufacturing) show different effects.
- Research Rabbit — starting from two or three central papers, explore the neighbourhood. Identify authors and clusters.
- Elicit — systematically work through the identified clusters. A table of 30 papers with columns for “study design, sample, country, effect direction”, so you have a structured overview when you start writing.
- SciSpace — dig deeper into the five to ten most important studies, clarify methodological details you missed on the first read.
And after all that: you read the papers you plan to cite, in full. No shortcut. No “I already saw that in Elicit”. The tool is the pre-selection — the citation is the original.
Where these tools reliably get it wrong
Four pitfalls to keep on your radar, no matter which tool:
Recency limit. All these tools sit on corpora that update with some lag. If your topic has moved significantly in the last six months — AI regulation, new social data — the very latest literature is missing.
Discipline bias. Semantic Scholar and friends cover the natural sciences, medicine, and computer science well. In the humanities and social sciences, the corpora are thinner. An Elicit run on a question in literary studies often returns only marginal hits.
Language. The corpora are English-heavy. German-language publications — especially in law, education, German literature — regularly fall through. If you want to reconstruct a debate in German, you have to supplement AI search with classic databases like wiso or GESIS.
Extraction errors. Even the best tool sometimes summarises a study wrong. A “70% pro” Consensus bar can hide a misclassification you only catch by opening the paper. Elicit will occasionally pull a number from the control group instead of the intervention arm. That’s not an excuse to skip the tool — it’s a reason to verify results before they end up in your thesis.
What no AI research tool checks
Finding and understanding papers is one side of the work. The other side only shows up at the end: whether every citation in your finished text actually carries the claim you hang on it. None of the four tools above checks that. Elicit tells you what the paper roughly contains. Consensus tells you whether the field roughly agrees. But whether the sentence on page 34 of your bachelor thesis actually reflects what’s on page 12 of the study — that’s outside their job description.
This is exactly where even careful work goes wrong: not out of bad intent, but because a claim drifts a little further from the source with every revision, or because the page number stops matching after reformatting. Plagiarism checkers don’t catch it either — they compare your text against other texts, not your claims against the sources.
That verification is what Acurio does. Not during research, but as the pass before you hand in:
- Upload your document — the thesis as a file, with your citations as they are.
- Automatic claim-versus-source check — for every citation, Acurio retrieves the underlying source, finds the passage you meant, and compares the claim in your text against what the source actually says, including the page number.
- A findings list per citation — for each entry you see whether the claim is supported, only partially covered, or not covered at all — and where, so you can fix it yourself.
To be clear about the limits: Acurio doesn’t judge whether your argument is good, and it doesn’t rewrite your text. It checks one thing — whether the source you cite supports the sentence you wrote.
If you want a first impression without an account: the Quick Check lets you paste a bibliography and flags the entries that need a closer look. A research tool helps you find the right source. The check before submission makes sure you use it correctly.
Bottom line
Elicit, SciSpace, Consensus, and Research Rabbit solve different problems. Using them all at once, high on possibility, mostly produces more noise. Using them in phases — Consensus for the overview, Research Rabbit for exploration, Elicit for structuring, SciSpace for understanding — saves you weeks of manual work.
But the core of scientific writing is the same as before the AI wave: you read the source you cite. You understand it. You represent it accurately. No tool takes that step off your desk. What tools can take off: the search before it, and the check after.
