Evidence Aggregation: Extracting Key Findings from Academic Papers (DEVIDENCE)
Policy Context
Over the past two decades, the body of evidence from academic impact evaluations has grown dramatically. For policymakers and donors committed to evidence-based decision-making, this means an ever-expanding pool of studies to draw on. The issue these decision-makers face is not a lack of evidence; rather, it is accessing that evidence and identifying the most effective interventions to improve welfare. This has sharpened interest in interventions that scale effectively and travel well across contexts. Aggregating findings from multiple studies into a unified database makes it easier to surface these interventions and draw meaningful comparisons across the literature. While meta-analyses remain the gold standard for evidence synthesis, many donors are shifting toward “living” evidence reviews—an approach that allows new findings to be continuously integrated as they emerge.
However, the existing infrastructure to produce evidence synthesis is labor intensive. On average, it takes a trained human coder 5–6 hours to extract data from a single study because key information is often scattered or locked in unstructured PDFs. Extracting data from hundreds of papers—each with different structures and reporting formats—requires immense time and labor. By using Large Language Models (LLMs), it’s possible to make it cheaper, faster, and potentially more accurate to extract this information. This project originated to test methods using LLMs for evidence aggregation, benchmarking the LLM extractions against human-coded datasets to continually ensure accuracy.
DEVIDENCE Design
BITSS and the Global Poverty Research Lab (GPRL) at Northwestern University are working together to test approaches to automate the extraction of key information from academic papers. We are leveraging a specialized metadata schema, co-developed with partners like the World Bank and J-PAL, to identify the information needed in academic papers for extracting treatment effects. By testing both fully automated and semi-automated approaches, the team is continually checking to see which workflows offer the best data with high accuracy and low cost.
DEVIDENCE leverages large language models (LLMs), deterministic code, and human-in-the-loop validation to construct and maintain our relational database. The integrated pipeline is designed to convert published studies and raw survey data into structured, comparable evidence, modeled on how researchers have traditionally done manual evidence aggregation. The pipeline is reproducible, documented, and designed to be run continuously as the literature grows:
Step 1: Metadata Collection and Verification
Publication metadata is collected through a combination of LLMs and API connections to external databases, then manually reviewed for accuracy.
Step 2: Structured Data Extraction
Treatment effects and other study-level fields are extracted from the full text of each paper using a combination of LLMs and deterministic code. No extracted value is trusted on first pass; coding scripts re-read each figure at its source, and a second model checks every field against the paper’s text. An AI agent orchestrator approves each stage of the extraction process before the pipeline proceeds to the next. Only once this multi-stage verification is complete does a record move on to human validation.
Step 3: Validation with DEVIDENCE Fellows
A cohort of scholars from low- and middle-income countries, convened through the DEVIDENCE Fellows Program, reviews and validates the extracted study-level information. Fellows work asynchronously through an application developed and maintained by the DEVIDENCE team.
Step 4: Cataloguing Microdata Variables
When underlying datasets are available, including household surveys, administrative records, and other individual-level datasets, the variables used in each analysis are systematically catalogued. This includes recording each variable’s original name, definition, units, and role within the study.
Step 5: Harmonizing Like Variables
Because studies often measure similar concepts using different definitions, scales, or naming conventions, all related variables are harmonized into a common framework using economic methods. This allows outcomes and covariates that are substantively comparable but reported under different labels or units to be aligned for cross-study comparison and synthesis.
Step 6: Quality Review
Every decision and judgment made by the LLMs is recorded, along with supporting evidence, creating a transparent audit trail for each piece of data. Before any information enters the DEVIDENCE database, humans review this audit trail and approve the AI’s work at key checkpoints, verifying the accuracy and quality of the aggregated data. The end result is a single database researchers can use, along with clear documentation of how each variable of interest was constructed.
Ultimately, BITSS and GPRL aim to build a publicly available dataset of treatment effects from global development papers, branded “DEVIDENCE” for Development Evidence. DEVIDENCE will serve as a resource for decision-makers interested in evidence for reducing poverty.
Results and Policy Lessons
Results forthcoming. The tool produced by this collaboration aims to enable the construction of large, open-access datasets of treatment effect data from social science papers. By using automation to speed up and reduce the cost of producing meta-analyses, DEVIDENCE will enable the global development community and other research disciplines to synthesize evidence, helping donors, researchers, and policymakers identify the most effective interventions based on academic research.