Consensus is an AI-powered academic search engine built to help researchers find, examine, and synthesize scholarly literature. You can enter a research question in ordinary language, and the platform searches its academic database before producing an answer supported by citations.
This order matters. Consensus retrieves papers first and then uses AI to interpret the material it finds. Its answers are linked to identifiable research rather than generated solely from a language model’s internal knowledge.
According to its documentation, Consensus searches more than 220 million peer-reviewed papers drawn from sources that include Semantic Scholar, OpenAlex, PubMed, and the company’s own crawl of the scholarly web. Its coverage extends across medicine, education, psychology, social science, engineering, and other research fields.
Consensus is particularly helpful during the early and middle stages of research: exploring a question, locating relevant papers, comparing findings, and identifying areas of agreement or disagreement. It can support a literature review, but it does not provide the reproducible search, screening, and appraisal process required for a formal systematic review.
Main Consensus AI Features
Three Search Modes
Consensus currently provides three levels of search:
- Paper Search returns a list of relevant studies without generating an AI synthesis.
- Pro examines up to 20 papers and produces a cited summary.
- Deep Review conducts a broader investigation across as many as 50 to 100 papers.
Deep Review breaks a question into sub-questions, runs up to 20 targeted searches, and may inspect more than 1,000 papers before selecting the studies used in its analysis. The finished report can include an introduction, methods, results, discussion, tables, visualizations, and citations. These are company descriptions of how the feature works, so the resulting report still needs to be checked against the original studies.

Consensus Meter
The Consensus Meter is one of the platform’s more recognizable features. It appears for suitable yes-or-no questions and classifies findings as Yes, No, Possibly, or Mixed.
The meter normally draws on between five and 20 of the most relevant results. Its detailed view provides information about publication recency, study methods, journal rankings, and citation counts for the papers associated with each position.
This is useful for obtaining an initial sense of where a body of research leans. It should not be interpreted as a scientific vote. Twenty studies are not necessarily representative of an entire field, and papers differ considerably in design, sample, context, and quality.
Advanced Research Filters
Search results can be filtered by:
- Publication date
- Open-access status
- Citation count
- Journal rank
- Study design
- Human or animal population
- Sample size
- Study duration
- Publisher
- Academic field
- Country
The study-design filters include randomized controlled trials, cohort studies, cross-sectional research, qualitative interviews, case studies, mixed-methods studies, systematic reviews, and meta-analyses.
Researchers can also describe some criteria in the query. For example:
Find studies published since 2022 on generative AI feedback in higher education. Focus on empirical studies involving undergraduate students.
Study Snapshot and Table View
Study Snapshot extracts details such as the research population, sample size, location, methods, outcomes, and results. Table View places selected attributes beside one another so that several studies can be compared.
Some attributes may be extracted only from an abstract. If the information is absent or unclear, the snapshot may remain incomplete. Researchers should return to the full paper before recording an extracted detail in an evidence matrix.
Research Agent and Follow-Up Threads
The Research Agent can manage multi-step queries. It may run several searches, follow citations, locate similar papers, compare studies, trace authors, and create structured tables. Researchers can continue with follow-up questions in the same thread without restating the original context.
Research Agent and conversational Threads are available through Pro and Deep searches.
Library, Full-Text Chat, and Zotero
Papers and searches can be saved in custom collections. Researchers can chat with individual papers, collections, uploaded documents, or their wider library when using Pro or Deep mode.
Consensus also supports Zotero imports. At present, the import is one-way, and it brings in the entire Zotero library rather than selected collections. Changes made in Consensus are not synchronized back to Zotero.
Search results and saved collections can be exported in CSV or RIS format for use in Zotero, Mendeley, EndNote, and other reference managers. Deep Review reports can be copied with citations or exported as PDF. Citation formats include APA, MLA, Chicago, Harvard, BibTeX, LaTeX, and numeric styles.
How to Use Consensus AI for Research
The most productive way to use Consensus is to move from broad exploration to close source checking. The following workflow uses an education research example:
How does generative AI feedback affect university students’ revision practices?
1. Test and Refine the Research Question
Begin with a broad Pro search to see how Consensus interprets the topic. The first results can reveal terms used by researchers that differ from the language in your original question.
For this example, relevant terms might include AI-generated feedback, automated feedback, revision behaviour, feedback uptake, and student engagement with feedback.
Run several versions of the question. A single wording can favour one set of papers and overlook another.
Possible searches include:
- How does generative AI feedback affect student revision?
- How do university students use automated feedback when revising academic writing?
- What factors influence student uptake of AI-generated writing feedback?
- Compare generative AI feedback with automated writing evaluation.
Record which terms consistently produce relevant results. This vocabulary can later inform searches in ERIC, Scopus, Web of Science, PsycINFO, or another disciplinary database.
2. Apply Criteria That Match the Question
Narrow the search using the population, research design, date range, and academic field.
For example:
Find empirical studies from 2022 onward examining generative AI feedback and writing revision among university students. Exclude commentaries and conceptual papers.
Do not apply study hierarchies mechanically. A randomized controlled trial may be useful for measuring an intervention’s effect, while interviews or classroom observations may be better for understanding why students accept, reject, or modify AI feedback.
3. Create an Initial Evidence Map
Use Table View or ask Pro to create a comparison with columns such as:
| Study | Participants | Context | AI tool or feedback type | Method | Revision outcome | Main limitation |
|---|
Treat this as a preliminary map. Check each row against the paper before transferring it into your research notes.
At this stage, look for patterns:
- Which student groups receive the most attention?
- How is revision quality measured?
- Are findings based on student perceptions or actual changes to drafts?
- Which countries and disciplines dominate the results?
- Are accessibility, language background, and prior writing ability considered?
These questions move the work beyond collecting summaries.
4. Investigate Agreement and Disagreement
A question such as “Does AI-generated feedback improve student writing?” may activate the Consensus Meter.
Inspect the studies assigned to each position. Differences in results may be explained by the kind of feedback, duration of the intervention, learner population, assessment criteria, or level of teacher involvement.
The meter is most helpful when it sends you back into the papers. If it reports mixed findings, ask:
Compare the methods and populations of studies reporting positive outcomes with those reporting limited or negative outcomes.
This gives you possible explanations to investigate. It does not establish why the results differ.
5. Examine the Most Relevant Papers Closely
Open Study Snapshot for promising papers, followed by the abstract and full text. Ask targeted questions:
- How was writing quality assessed?
- Who evaluated the revisions?
- Did the study distinguish between accepting feedback and learning from it?
- Was there a comparison group?
- How long did the intervention last?
- What limitations did the authors report?
Save the papers into a focused collection. I would also export them to Zotero or another reference manager rather than allowing Consensus to become the only location for the project.
6. Use Deep Review to Challenge Your Coverage
Once you understand the topic, run a Deep Review with explicit instructions:
Review empirical research on generative AI feedback and university writing revision since 2022. Compare research designs, populations, feedback types, learning outcomes, and limitations. Include contradictory findings and identify populations that are underrepresented.
Compare the resulting report with your existing evidence map. The real value is in spotting papers, terms, or disagreements you may have missed. The generated literature review should remain a research aid, not text to submit as your own review.
7. Verify and Document the Search
Open every source that supports an important claim. Check whether the synthesis accurately represents the population, method, findings, and qualifications in the paper.
For a narrative review, record the search date, exact queries, filters, and databases consulted. For a systematic or scoping review, Consensus can help with preliminary exploration and supplementary searching, but it should not replace a registered protocol, database-specific strategies, duplicate screening, quality appraisal, and a documented PRISMA-style selection process.
Consensus AI Pricing
Pricing checked in September 2026:
| Plan | Price | Main allowance |
|---|---|---|
| Free | $0 | Unlimited paper searches, 10 Pro messages, 3 Deep Reviews and 10 Study Snapshots per month |
| Pro | $20 monthly or $144 annually | Unlimited Pro messages, 15 Deep Reviews and unlimited Study Snapshots |
| Deep | $65 monthly or $540 annually | 200 Deep Reviews, unlimited Pro messages and unlimited Study Snapshots |
| Teams | Custom | Shared administration, centralized billing and 50 Deep Reviews per user |
| Enterprise | Custom quote | Institutional management, library integrations and higher-volume access |
Students and faculty can currently apply for a 40% discount, while clinicians may qualify for 25% off.
Consensus Compared with Other Research Tools
| Tool | Best use | Important distinction |
|---|---|---|
| Consensus | Asking research questions and obtaining cited evidence summaries | Particularly accessible for question-based searching |
| Elicit | Screening papers and building structured extraction tables | Better suited to systematic evidence extraction |
| Scite | Examining how later papers cite a study | Stronger for citation context and claim checking |
| SciSpace | Reading and questioning individual papers | More focused on PDF analysis |
| ResearchRabbit | Exploring related papers, authors, and citation networks | Better for visual discovery |
| Zotero | Managing references, PDFs, notes, and bibliographies | A reference manager rather than an evidence-synthesis engine |
Limitations and Responsible Use
Consensus searches a very large database, but it does not contain every study. Its results should be supplemented with disciplinary databases and library searches when coverage matters.
The AI can also misread a real paper. Consensus itself identifies source misinterpretation as a remaining risk, even though its retrieval-first design reduces fabricated citations.
Researchers should also remember that citation count, journal rank, and recency are indicators rather than measures of truth. Highly cited studies can contain weaknesses, while new or less-cited work may be methodologically sound.
Consensus states that prompts, uploaded content, and personal information are not used to train its own or third-party language models. However, submitted material may be processed by service providers including OpenAI, Anthropic, Google, and Baseten to produce requested results. Avoid uploading confidential participant data, unpublished peer-review material, embargoed manuscripts, or sensitive institutional documents without permission.
Final Assessment
Consensus is a useful starting point when a researcher wants to ask a direct question and see which studies speak to it. The filters, study comparisons, Consensus Meter, full-text chat, Research Agent, and Deep Review provide several routes into the literature.
Its outputs become useful research only after the underlying papers have been read and evaluated. Used in that way, Consensus can shorten the distance between an initial question and a defensible collection of sources without taking over the interpretive work that belongs to the researcher.








