hackforla / hackforla/data-science

Decline in US R&D productivity?

Open
#279 0 comments 0 reactions 0 assignees View on GitHub
complexity: Large CoP: Data Science project duration: one time role: data analysis role: data science size: 8pt
Dominant language
Jupyter Notebook
Stars
33
Forks
22
PR merge metrics
No merged PRs in 30d

Description

# Investigate Whether U.S. Research Productivity Is Declining

## Overview

The United States spends more on research and development and employs more researchers than it did several decades ago, yet economic productivity growth, disruptive scientific discoveries, and the translation of research into broad social benefits may not have increased proportionally.

A prominent research hypothesis argues that ideas are becoming harder to find. In fields such as semiconductors, agriculture, medicine, and industrial innovation, maintaining a similar rate of improvement has often required substantially more researchers and investment. One widely cited study estimated that sustaining Moore’s Law required more than 18 times as many researchers as in the early 1970s.

However, this conclusion is not settled. A 2026 firm-level study found that patents per unit of R&D input were increasing or stable and suggested that the larger problem may be that innovations are producing less subsequent business growth. A separate 2026 review challenged both the conceptual model and empirical measures behind claims of sharply declining research productivity.

The issue is therefore not simply whether the United States produces enough papers or patents. It is whether the research system efficiently produces:

* Important scientific discoveries;
* Reliable and reproducible knowledge;
* Useful technologies;
* New companies and industries;
* Improvements in health, energy, manufacturing, and living standards;
* A sustainable scientific workforce;
* Long-term research capacity.

The central research question is:

> **Is U.S. research becoming less productive, and which funding structures, institutional practices, team designs, and incentives produce the greatest scientific and social value per unit of research effort?**

The investigation should distinguish:

* **Publication productivity:** papers per researcher or research dollar;
* **Discovery productivity:** important advances per researcher or dollar;
* **Translation productivity:** movement from research to practical application;
* **Economic productivity:** effect of research on firm growth, industry productivity, and new markets;
* **Social productivity:** health, environmental, safety, or public benefits;
* **Workforce productivity:** sustainable output without excessive attrition or administrative burden;
* **Resilience:** the ability of the research system to continue functioning through funding or institutional shocks.

The recommended units of analysis are:

* Researcher-year;
* Grant-year;
* Institution-field-year;
* Research-team or publication;
* Firm-year;
* Scientific field-year.

---

## Action Items

### Initial Evidence

Several findings justify the investigation:

* U.S. R&D expenditures have continued to increase, and business now performs most U.S. development and a large share of applied research. NCSES provides annual expenditure data by sector, source, and type of research through 2024.
* Earlier studies found that research effort rose rapidly while measured research productivity declined in semiconductors, agriculture, medicine, and firm-level innovation.
* Recent research disputes a universal decline. A 45-year panel of U.S. firms found no systematic secular decline in patent production after accounting for R&D inputs. It instead found that firm growth conditional on patent growth had weakened.
* Large research teams tend to develop and extend established ideas, while smaller teams have often been associated with more disruptive work. This raises questions about whether team growth changes the type rather than the quantity of discovery.
* Publication and citation counts have grown, but these metrics may be inflated by larger teams, increased international collaboration, publication pressure, field growth, and citation conventions.
* U.S. federal research support experienced substantial disruption in 2025. Nature reported that NSF awarded 25% fewer new grants than its prior ten-year average, while NIH awarded 24% fewer. This shock may provide evidence about how funding stability affects laboratories, careers, and output.
* A 2026 survey of nearly 1,000 NIH-funded researchers found that more than one-quarter had laid off laboratory members, more than two-fifths had canceled planned research, and two-thirds had advised students to consider careers outside academia. The survey is not representative of all U.S. science, but it indicates severe disruption in affected laboratories.
* Researcher hiring may not always promote innovation. A 2025 working paper proposed that incumbent firms sometimes hire researchers defensively, reducing entry and creative destruction despite greater R&D spending.
* China and other countries have increased their scientific output and capacity. NCSES provides field-level publication data by country, including chemistry, astronomy, engineering, and other disciplines.

These findings establish a plausible productivity problem but do not establish that all fields are declining or that fewer resources would improve science.

---

### Hypotheses

#### H1 — Important discoveries require increasing research effort

As the stock of knowledge expands, researchers may need more training, equipment, data, and specialization to reach the scientific frontier.

**Evidence to seek:**

* Research teams and expenditures rise faster than breakthrough output;
* Time from training to independent contribution increases;
* Discovery costs rise within stable research objectives;
* More narrow specialization is needed for comparable advances;
* The pattern differs between mature and emerging fields.

#### H2 — Conventional metrics reward quantity over importance

Publication counts, journal placement, citations, and grant totals may encourage incremental work, fragmented papers, and fashionable topics.

**Evidence to seek:**

* Higher paper counts do not predict more disruptive or reproducible work;
* Evaluation incentives increase publication volume without increasing practical outcomes;
* Researchers alter topic selection around funding priorities;
* Institutions rewarded for volume produce fewer high-impact discoveries per dollar;
* Metric reforms change researcher behavior.

#### H3 — Larger teams favor development while smaller teams favor disruption

Large collaborations may excel at validating, scaling, and extending ideas, whereas small teams may explore less established directions.

**Evidence to seek:**

* Team size predicts citation structure and novelty;
* Small-team discoveries are more likely to introduce new combinations;
* Large teams are more likely to consolidate established fields;
* Effects persist after controlling for discipline and project scale;
* Funding portfolios containing both team types outperform homogeneous portfolios.

#### H4 — Administrative burden reduces time available for research

Grant writing, compliance, reporting, hiring, procurement, and institutional administration may consume an increasing portion of researchers’ time.

**Evidence to seek:**

* Researchers spend fewer hours on direct research;
* Grant success rates decline while application effort increases;
* Administrative staffing and costs rise faster than research output;
* Simplified grant programs produce similar quality at lower overhead;
* Early-career scientists experience the greatest burden.

#### H5 — Funding has become too concentrated

A small group of institutions, laboratories, senior researchers, or firms may receive a growing share of resources.

**Evidence to seek:**

* Grant concentration rises over time;
* Marginal scientific output declines at the highest funding levels;
* Smaller awards to more researchers produce greater portfolio-level novelty;
* Early-career investigators have lower success rates after controlling for proposal quality;
* Geographic or institutional concentration reduces topic diversity.

#### H6 — Short grant cycles discourage high-risk and long-horizon research

Researchers dependent on frequent renewals may avoid projects with uncertain or delayed results.

**Evidence to seek:**

* Longer grants support more novel or disruptive research;
* Short funding horizons increase incremental publication strategies;
* Funding interruptions cause irreversible team and data losses;
* Stable institutional support improves long-term outcomes;
* Effects vary between theoretical, experimental, and infrastructure-intensive fields.

#### H7 — Research productivity is weakened by poor reproducibility

Findings that cannot be reproduced consume resources, misdirect later work, and weaken public trust.

**Evidence to seek:**

* Fields with low replication rates show slower cumulative progress;
* Data and code availability predict reproducibility;
* Independent replication prevents costly downstream failure;
* Incentives for null results reduce publication bias;
* Replication spending has positive long-term value despite producing fewer novel papers.

#### H8 — Translation from science to economic growth has weakened

The research system may continue generating ideas while firms and institutions struggle to convert them into productivity, businesses, or scaled technologies.

**Evidence to seek:**

* Patent output remains strong while productivity growth weakens;
* Time from patent to commercialization increases;
* New firms capture a smaller share of innovation;
* Market concentration reduces diffusion and creative destruction;
* Manufacturing, regulatory, capital, or workforce constraints obstruct scale-up.

This hypothesis is consistent with 2026 findings that patent productivity may be stable while growth conditional on patenting has declined.

#### H9 — Defensive corporate R&D reduces genuine innovation

Large incumbent firms may hire scarce researchers partly to prevent competitors from using their talent.

**Evidence to seek:**

* Researcher hiring by dominant firms reduces startup entry;
* Patents rise without proportional productivity growth;
* Research labor is concentrated in incumbents;
* Acquisitions reduce continuation of competing research programs;
* Effects are strongest in concentrated industries.

#### H10 — AI can increase research productivity, but may favor incremental work

AI can accelerate literature review, coding, simulation, analysis, and experiment design, but it may also encourage research near dense existing knowledge.

**Evidence to seek:**

* AI adoption reduces time to publication or experiment completion;
* Results differ between exploratory and incremental research;
* AI-generated hypotheses are more or less novel than human-generated hypotheses;
* Productivity gains are concentrated among specific disciplines;
* Error, reproducibility, and review burdens change with AI adoption.

A 2025 theoretical analysis argues that AI’s effect may depend on whether it primarily interpolates within existing knowledge or enables exploration beyond it.

#### H11 — Funding instability damages productivity beyond the value of the funding removed

Sudden cuts may disperse teams, terminate experiments, interrupt longitudinal datasets, and drive researchers out of science.

**Evidence to seek:**

* Interrupted grants produce disproportionate publication and career losses;
* Restarting a laboratory costs more than maintaining continuity;
* Early-career researchers have higher exit rates after funding shocks;
* Collaborations weaken after sudden national funding reductions;
* Effects persist after funding is restored.

#### H12 — Research productivity differs sharply by field and output definition

There may be no valid single national productivity measure.

**Evidence to seek:**

* Clinical, theoretical, engineering, and observational fields show different trends;
* Papers, patents, discoveries, products, and health gains produce conflicting rankings;
* Capital-intensive fields appear less productive under labor-only metrics;
* Field-normalized measures change conclusions;
* Some fields improve while others stagnate.

---

### Research Plan

#### 1. Define multiple research-output measures

Do not treat papers or patents as the final outcome.

Measure:

* Publications;
* Field-normalized citations;
* Novelty;
* Disruption;
* Replications;
* Retractions and corrections;
* Patents;
* Patent citations;
* Clinical or technical adoption;
* New products;
* Startup formation;
* Productivity growth;
* Health or environmental improvements;
* Standards, datasets, and open-source software.

Research productivity may be expressed as:

Image

Inputs should include:

* Expenditure;
* Researchers;
* researcher hours;
* equipment;
* computing;
* facilities;
* administrative effort;
* elapsed time.

#### 2. Build a field-year panel

Recommended period:

* 1980 through the latest available year;
* 2000 onward for detailed author-, grant-, and institution-level analysis.

Fields should include:

* Physics;
* chemistry;
* materials science;
* biology and medicine;
* computer science;
* engineering;
* energy;
* agriculture;
* earth and environmental science;
* social science as a comparison.

Core variables:

* R&D spending;
* number of researchers;
* grants;
* publications;
* citations;
* patents;
* team size;
* collaboration;
* investigator age;
* institution;
* funding concentration;
* economic and social outcomes.

#### 3. Construct alternative productivity indexes

Compare at least four definitions:

Image

Image

Image

Image

The investigation should report when these measures disagree.

#### 4. Measure research novelty and disruption

Use citation-network measures to determine whether a publication:

* Develops prior work;
* Combines previously disconnected areas;
* Displaces earlier approaches;
* Creates a new research trajectory.

Do not label low-disruption research as low value. Verification, measurement, infrastructure, and incremental improvement may be essential.

#### 5. Analyze team size

For each research output, record:

* Number of authors;
* institutions;
* countries;
* disciplines;
* principal investigators;
* funding sources.

Model:

Image

Possible outcomes:

* Novelty;
* disruption;
* citations;
* replication;
* patenting;
* commercialization;
* time to completion.

The nonlinear term tests whether team size has diminishing or changing effects.

#### 6. Measure funding concentration

Calculate:

* Share of funding held by the top 1%, 5%, and 10% of investigators;
* Institutional concentration;
* Geographic concentration;
* Field concentration;
* Herfindahl–Hirschman Index.

Test whether output per marginal dollar declines as investigator or institution funding increases.

#### 7. Study funding shocks

Use grant interruptions, policy changes, budget reductions, or delayed appropriations as natural experiments.

Compare affected and less-affected:

* Fields;
* institutions;
* laboratories;
* investigators;
* cohorts.

Outcomes should include:

* Laboratory employment;
* publications;
* collaboration;
* graduate admissions;
* career exits;
* patenting;
* future funding.

The 2025 reductions in new NSF and NIH awards offer a recent shock, but the analysis must distinguish temporary administrative delays, multiyear award accounting, canceled grants, and durable funding reductions.

#### 8. Analyze career-stage effects

Measure separately:

* Graduate students;
* postdoctoral researchers;
* early-career faculty;
* mid-career researchers;
* senior investigators;
* research staff and technicians.

Test whether:

* Early-career researchers receive fewer independent opportunities;
* Funding shocks disproportionately remove younger scientists;
* Time to independence has increased;
* Career insecurity changes research risk-taking;
* Attrition reduces future output beyond immediate grant losses.

#### 9. Estimate administrative burden

Potential measures:

* Number of proposals submitted per award;
* application length;
* review duration;
* reporting requirements;
* indirect-cost rates;
* administrative staff;
* procurement time;
* researcher survey estimates of time allocation.

Calculate:

Image

This estimates the total effort required across successful and unsuccessful applications to fund one project.

#### 10. Examine reproducibility and reliability

Track:

* Retractions;
* corrections;
* replication studies;
* data availability;
* code availability;
* preregistration;
* sample sizes;
* statistical power;
* result consistency.

Compare output systems that reward:

* Publication volume;
* open science;
* replication;
* shared infrastructure;
* negative results;
* long-term datasets.

#### 11. Analyze translation into economic and social outcomes

Construct pathways such as:

Image

Estimate:

* Time between stages;
* probability of reaching each stage;
* institutional and field differences;
* failure points;
* effect of manufacturing and market structure.

Avoid assuming every valuable scientific result should become a patent or company.

#### 12. Evaluate AI-assisted science

Identify early-adopting laboratories or fields and compare changes in:

* Research cycle time;
* output;
* novelty;
* error;
* team composition;
* replication;
* compute cost.

Use staggered adoption designs where credible. Avoid attributing general post-2022 changes entirely to generative AI.

#### 13. Identify positive-deviant research systems

Find fields, institutions, programs, or funders that produce unusually strong outcomes given:

* Funding;
* researcher count;
* field maturity;
* equipment needs;
* institution size.

Study practices including:

* Longer grants;
* smaller grants distributed broadly;
* high-risk portfolios;
* investigator autonomy;
* shared facilities;
* open data;
* interdisciplinary teams;
* replication funding.

#### 14. Conduct robustness checks

Repeat results using:

* Papers;
* citations;
* patents;
* disruption;
* commercialization;
* societal outcomes;
* researcher count;
* researcher hours;
* real spending;
* different citation windows;
* field-normalized metrics;
* fractional and whole counting;
* domestic and international collaborations.

#### 15. Produce the Final PowerPoint

Recommended structure:

1. What research productivity means;
2. Rising research inputs;
3. Papers and patents over time;
4. Are ideas becoming harder to find?;
5. Evidence challenging the decline hypothesis;
6. Team size and research type;
7. Funding concentration;
8. Administrative and career burden;
9. Reproducibility;
10. Translation into economic and social value;
11. Funding shocks and system resilience;
12. Policy options and evaluation metrics.

The presentation must distinguish:

* Output from value;
* Papers from discoveries;
* Patents from adoption;
* Citations from reliability;
* Research productivity from economic productivity;
* Basic research from commercialization;
* Incremental science from low-value science;
* Short-term output from long-term capacity.

---

## Resources

### Data, APIs, and Where to Get Data

#### OpenAlex API

OpenAlex provides open records for:

* Publications;
* authors;
* institutions;
* topics;
* citations;
* funders;
* grants;
* affiliations;
* countries.

Use for:

* Publication output;
* citation networks;
* team size;
* institutional mobility;
* international collaboration;
* field classification;
* novelty and disruption measures.

OpenAlex should be the principal open bibliometric source.

#### Crossref API

Use for:

* DOI metadata;
* publication dates;
* references;
* funder information;
* licenses;
* updates and corrections.

Crossref can validate publication metadata and supplement OpenAlex.

#### Semantic Scholar API

Use for:

* Papers;
* citations;
* authors;
* fields;
* reference networks;
* selected influence measures.

Coverage and definitions should be compared with OpenAlex before combining records.

#### ORCID Public Data

Use for:

* Researcher employment histories;
* education;
* publications;
* grants;
* persistent researcher identifiers.

Participation is voluntary and incomplete, so ORCID should supplement rather than define the population.

#### NSF Award Search API

Use for:

* Awards;
* investigators;
* institutions;
* programs;
* directorates;
* award amounts;
* dates;
* abstracts.

Link grants to publications, patents, investigators, and subsequent funding where possible.

#### NIH RePORTER API

Use for:

* NIH projects;
* investigators;
* organizations;
* funding;
* publications;
* patents;
* clinical studies;
* project terms.

It is especially useful for developing and validating grant-to-output pipelines.

#### USAspending API

Use for:

* Federal research awards;
* obligations;
* recipients;
* agencies;
* transactions;
* geographic distribution.

This supports cross-agency comparisons beyond NSF and NIH.

#### NCSES and Science and Engineering Indicators

NCSES provides downloadable tables for:

* U.S. R&D expenditures;
* research performers;
* funding sources;
* publication output;
* international comparisons;
* science and engineering workforce;
* field-level publications.

The 2025 discovery report provides downloadable R&D and publication tables, while the 2026 State of U.S. Science and Engineering includes updated R&D expenditure data through 2024.

#### Higher Education Research and Development Survey

Use for:

* University R&D expenditures;
* fields;
* funding sources;
* institution;
* research equipment;
* federally financed R&D.

#### Survey of Federal Funds for Research and Development

Use for:

* Federal agency obligations;
* fields;
* performers;
* basic and applied research;
* development.

#### Business Enterprise Research and Development Survey

Use for:

* Business R&D;
* industry;
* company size;
* funding;
* employment;
* domestic and international activity.

Detailed microdata may require restricted access, but public aggregates support field- and industry-level analysis.

#### USPTO Open Data

Use for:

* Patents;
* inventors;
* assignees;
* citations;
* classifications;
* prosecution;
* ownership.

Link patents with publications, grants, institutions, and firms.

#### PatentsView Historical Data

Use historical bulk files where compatible with current USPTO data.

Verify current endpoint and migration status before building production workflows.

#### SEC EDGAR APIs

Use for:

* Firm R&D expenditure;
* financial performance;
* acquisitions;
* employment where reported;
* commercialization indicators.

Firm identities must be carefully matched to patent assignees and subsidiaries.

#### Census Business Dynamics Statistics

Use for:

* Startup formation;
* establishment entry and exit;
* job creation;
* firm age;
* industry dynamics.

#### Bureau of Labor Statistics APIs

Use for:

* Research occupations;
* wages;
* employment;
* labor productivity;
* industry trends.

#### ClinicalTrials.gov API

Use where appropriate for:

* Study registration;
* completion;
* results reporting;
* sponsors;
* interventions;
* development timelines.

#### Federal RePORTER Archives and Agency Data

Use to extend historical grant coverage across federal agencies where available.

#### Retraction Watch Database

Use for:

* Retractions;
* reasons;
* fields;
* institutions;
* journals.

Confirm current licensing and access requirements before redistribution.

#### Open Science Framework API

Use for:

* Preregistrations;
* open materials;
* data;
* project histories;
* replications.

#### GitHub API

Use cautiously as a measure of research software and open-source output:

* Repositories;
* releases;
* contributors;
* dependencies;
* citations where linked.

Repository activity is not directly comparable across fields.

### Foundational Resources

* Bloom, Jones, Van Reenen, and Webb, *Are Ideas Getting Harder to Find?*
* Fort, Goldschlag, Liang, Schott, and Zolas, *Growth Is Getting Harder to Find, Not Ideas*.
* *Research-Driven Productivity Growth Redux: Are Ideas Really Getting Harder to Find?*
* Fernández-Villaverde, Yu, and Zanetti, *Defensive Hiring and Creative Destruction*.
* NCSES, *Discovery: R&D Activity and Research Publications*.
* Nature, *U.S. Science After a Year of Trump: What Has Been Lost and What Remains*.

### Central Scientific Caution

Research productivity cannot be reduced to:

Image

That ratio can increase when papers become shorter, teams divide results into more publications, or citation practices change. It can decline when research becomes more rigorous, capital-intensive, collaborative, or focused on difficult problems.

The strongest investigation will triangulate:

1. Research inputs;
2. Publication output;
3. Novelty and disruption;
4. Reliability and replication;
5. Patents and technical outputs;
6. Adoption;
7. Economic and social impact;
8. Workforce sustainability;
9. Time required;
10. System resilience.

The most important finding may be that U.S. science is still producing ideas efficiently, but the country has become less effective at **diffusing, scaling, and converting those ideas into broad productivity and public benefit**.

- If this issue requires access to 311 data, please answer the following questions:
- Do you need a one-time or ongoing dump of the data?
- Do you need subset of data (i.e. certain years) or the entire data set (approx. 4 million rows or 11 GB)?
- If a subset is needed, please define subset characteristics (i.e. date range, etc.)
- Do you need online access via an API or a download of data?

Contributor guide

Open the contributing guide

Research direction

No files, tests, or entry points are identified in the issue. First define a bounded field, dataset, metric, and reproducible analysis plan from the listed hypotheses; done would require a documented investigation with stated evidence and limitations.

Written by the indexing model from the issue text.

Assessment

Tech stack
jupyter-notebook
Domain
analytics, data
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.