

Elicit is an AI research assistant that searches 138 million academic papers by meaning, not just keywords. It extracts structured data from studies into tables, automates literature screening, and exports to Zotero. Primarily used by PhD students and systematic review teams.
Elicit is an AI research assistant built specifically for academic literature search and data extraction. It was developed at Ought, a non-profit AI safety research organization founded by Andreas Stuhlmüller, a former Stanford Computation and Cognition Lab researcher, and spun out as an independent public benefit corporation in September 2023. Elicit's core problem: academic researchers spend an average of four weeks manually screening papers for a single systematic review. Elicit cuts that screening window from weeks to hours by running semantic, natural-language queries across a database of 138 million papers and 545,000 clinical trials.
The tool's primary value is structured extraction: researchers add custom columns to pull sample sizes, intervention descriptions, outcome measures, and methodology notes from dozens of papers simultaneously. Every extracted data point links back to a highlighted passage in the source paper, which makes human verification faster. Elicit also generates Automated Research Reports, exports tables to CSV, BibTeX, and RIS formats for direct import into Zotero, Mendeley, or EndNote, and supports PDF upload for chat-style question answering against private document libraries. A High-Accuracy Mode, available on paid tiers, uses advanced LLM backends for complex extraction tasks. As of February 2025, Elicit reported 400,000 active monthly researchers after raising a $22M Series A at a $100M valuation.
What Elicit actually is in April 2026
Elicit is not a general-purpose AI chatbot with a research skin. It is a structured literature workflow tool. The distinction matters: when you enter a query, Elicit does not generate prose from its own knowledge. It searches OpenAlex and Semantic Scholar's indexed corpus, ranks papers by semantic relevance to your question, and extracts named fields from abstracts and full text where available.
The flagship feature is the extraction table. You define columns, either from a library of 35 pre-built options (population, intervention, sample size, study design, follow-up period, effect size) or via custom natural-language prompts ("What statistical method did this paper use for confounding?"), and Elicit populates each row with a cited text snippet from the corresponding paper. The result is a structured matrix that can be exported and analyzed, similar to what a research assistant would produce after reading 40 papers, but generated in minutes rather than days.
The Systematic Review workflow goes further: it adds a screening phase (include/exclude with AI-suggested reasoning), generates a PRISMA-compatible flow diagram, and produces a summary report. This pipeline is what draws clinical researchers, public health teams, and evidence synthesis groups to the platform. Elicit's February 2025 Series A announcement named healthcare, policy, and pharma as core expansion markets beyond academia.
Where Elicit sits versus Consensus and SciSpace
Three tools dominate AI-assisted academic research as of April 2026. They share Semantic Scholar's underlying paper corpus but pursue different workflows with meaningfully different mechanics.
Consensus is built for rapid claim verification rather than systematic data collection. Its core mechanic is a binary "Yes/No Consensus Meter" that aggregates findings from relevant studies to answer a single research question in seconds. Consensus also surfaces SJR (SCImago Journal Ranking) scores and SciScore indicators alongside results, giving researchers a built-in signal about journal quality that Elicit lacks entirely. Where Elicit builds multi-column extraction tables across dozens of papers, Consensus gives a quick verdict on a single scientific claim. Consensus costs $20/month at entry versus Elicit's $12/month, and neither tool supports languages other than English.
SciSpace takes a PDF-first approach. Its core strength is explaining individual complex papers: upload a paper and ask the AI to walk through an equation, interpret a statistical table, or summarize a dense methods section in plain language. SciSpace supports up to 50 custom extraction columns on paid plans versus Elicit's maximum of 10, includes a built-in writing assistant (Elicit has no writing support at all), and provides a dedicated citation generator. Its $20/month entry price is higher than Elicit's, and its free plan is more restrictive than Elicit's in some respects. SciSpace's modular interface separates tools into panels, while Elicit's interface is minimalist and linear. The mechanical choice comes down to workflow priority: use SciSpace when you need to understand individual papers deeply; use Elicit when you need consistent data extracted across a large collection for a formal review.
"Elicit was not developed to provide you with deep analysis and synthesis. It may miss nuances or complex relationships between studies that a human expert would catch. But it's much easier to verify information than to collect it in the first place." - Boris Nikolaev, researcher, Medium, 2025
How AI actually works inside Elicit
Elicit's search layer uses vector embeddings to match query intent against paper abstracts rather than keyword co-occurrence. This is the reason Elicit finds relevant papers that use different terminology than your query terms, which is essential in research fields where the same concept carries different names across disciplines or time periods.
Extraction columns run a separate inference pass over each paper. For each paper-column pair, Elicit sends the relevant paper sections to a language model backend along with your column definition. The model extracts the most relevant snippet, and Elicit stores the source location. In High-Accuracy Mode (paid tiers only), Elicit routes requests to more capable LLM backends for columns requiring complex interpretation. The credit system reflects this: simple extraction columns are cheap, high-accuracy columns burn significantly more credits per paper.
PDF uploads add a retrieval-augmented generation layer. When you chat with an uploaded paper, Elicit retrieves the most relevant passages to your question and passes them to the LLM with instructions to answer from the document only. Source attribution is shown for every response. This is why Elicit's citation accuracy is substantially higher than general-purpose AI tools for paper-specific questions, though it still requires human verification for any claim you plan to publish.
The accuracy and reproducibility concerns researchers keep raising
Elicit's marketing claims 80-90% extraction accuracy. Academic studies published in 2025 paint a more nuanced picture that every researcher should understand before relying on the tool for formal work.
A June 2025 comparative study published in Cochrane Evidence Synthesis and Methods (Bianchi et al.) analyzed Elicit's extraction against human reviewers across 20 randomized controlled trials. The findings: across all seven variables studied, Elicit extracted "equal" information to human reviewers in only 20.7% of cases, was "partially equal" (missing information) in 45.7%, and actively "deviating" (incorrect data) in 4.3%. Critically, 95% of "intervention effects" extractions, the single most important variable in clinical research, were rated only partially equal. The study concluded: "Verification by human reviewers is necessary to ensure that all relevant information is captured completely and correctly by Elicit."
A separate Cochrane study (Lau et al., 2025) found that Elicit's average sensitivity in literature searches was approximately 39.5%, meaning the tool misses roughly 60% of potentially relevant papers in a comprehensive search. Elicit's precision was higher at 41.8%, but sensitivity is what matters for systematic reviews. The authors concluded that Elicit "did not search with high enough sensitivity to replace traditional literature searching."
"Elicit should serve as a research scaffold, not an authoritative source. Users require instruction on interrogating AI outputs and understanding transparency limitations, particularly regarding data provenance and search methodology documentation." - Deakin University Library, AI Tool Evaluations guide, 2025
Additional documented limitations: Elicit cannot extract information from figures, flowcharts, or visual data tables within PDFs. Searches are not reproducible in the structured, documented way required for PRISMA-compliant systematic review reporting. The tool is English-only, restricting usefulness in multilingual research contexts.
These are not reasons to avoid Elicit. They are reasons to use it correctly: as a screening and acceleration layer, not as a replacement for human review. Researchers who understand the tool's actual performance profile consistently get value from it; researchers who take outputs at face value create downstream problems.
Who Elicit is for
Graduate students and PhD researchers in empirical fields are Elicit's primary user base. If your research involves structured data (clinical trials, experimental studies, surveys), Elicit's extraction tables are genuinely transformative for first-pass screening and data collection. Biomedical, psychology, social science, and machine learning researchers get the most value. Elicit's semantic search is particularly good in fields where terminology evolves across decades of literature.
Systematic review and evidence synthesis teams in clinical research, public health, and policy use Elicit's Pro and Team tiers for the full systematic review pipeline: screening, extraction, and report generation. The ability to export PRISMA-compatible documentation (with human verification built into the workflow) makes it useful even given the sensitivity limitations.
Skip Elicit when you work primarily in humanities, law, or qualitative social science, where papers use narrative reasoning rather than structured data and Elicit's extraction tables add little value. Skip it if you need comprehensive, reproducible search coverage for a formal publication-ready systematic review without supplementing with traditional database searches in PubMed, Scopus, or Web of Science. Skip it if you are a casual user who wants quick science-backed answers to single questions, where Consensus is a better fit. Skip it if your research is in a language other than English.
The free tier's two-report-per-month cap means most researchers who try it hit the ceiling within one afternoon and need to decide on a paid plan. The $10/month annual Plus rate is where most individual researchers land. For teams running multiple systematic reviews per month, the Pro tier at $42/month annually becomes the appropriate entry point.
User Reviews
No reviews yet. Be the first to share your experience!
Sign in to write a review.
Related articles
Guides and articles related to Elicit.

Perplexity AI Tutorial: Maximize Your Research Workflows

Gemini Deep Research for a Thesis Literature Review in One Day (2026)

The Personal AI Productivity Stack (2026): One Tool Per Job, Nothing Extra

TAM SAM SOM Market Sizing With Perplexity: The Sourcing Pattern That Survives Pushback (2026)

The Perplexity Due-Diligence Prompt Library: 10 Prompts That Surface Red Flags (2026)
