Papers That Shaped My Thinking
Papers that have shaped how I think about causal inference, urban environments, behavioral economics, and computational modeling — and that I recommend and teach in my classes.
▾Urban Economics10 papers
Landmark Papers
Moving to Opportunity for Fair Housing
Randomized housing voucher experiment — neighborhood environments shape long-run economic and health outcomes.
Why it's a landmark: First to randomize families into lower-poverty neighborhoods, setting the causal benchmark that reframed the entire neighborhood-effects literature.
2005
Why Have Housing Prices Gone Up?
Land use regulations — zoning restrictions drive up housing costs more than construction costs.
Why it's a landmark: Decomposed house prices into construction cost, land, and a regulatory 'zoning tax,' attributing high coastal prices to supply regulation rather than scarcity.
2005
Estimates of the Impact of Crime Risk on Property Values from Megan's Law
Housing markets respond sharply to newly salient crime information — direct motivation for salience-based housing research.
Why it's a landmark: Used sex-offender arrival and departure as a localized natural experiment to estimate a sharp house-price gradient in crime risk.
2008
The Fundamental Law of Road Congestion
Adding roads increases driving proportionally — transportation infrastructure and urban sprawl.
Why it's a landmark: Showed vehicle-kilometers traveled rise one-for-one with road capacity, establishing induced demand so new roads do not relieve congestion.
2011
Local Economic Development, Agglomeration Economies, and the Big Push
Tennessee Valley Authority — place-based industrial policy raised manufacturing wages but had mixed aggregate effects.
Why it's a landmark: Used the TVA to estimate agglomeration spillovers and evaluate big-push place-based policy, finding lasting manufacturing gains but modest aggregate efficiency effects.
2014
The Effects of Exposure to Better Neighborhoods on Children
Uses families' moves across areas to study how childhood neighborhood exposure affects long-run adult outcomes.
Why it's a landmark: Used movers and sibling age-at-move variation to show neighborhood exposure effects on children accumulate roughly linearly with childhood years spent there.
2018
Housing Constraints and Spatial Misallocation
Studies how housing supply constraints in high-productivity cities affect the spatial allocation of workers and aggregate output.
Why it's a landmark: Quantified aggregate output lost to housing-supply restrictions in high-productivity cities, estimating large national gains from relaxing them.
2019
Creating Moves to Opportunity: Experimental Evidence on Barriers to Neighborhood Choice
Uses a randomized housing-voucher experiment to examine how search assistance and information barriers shape low-income families' neighborhood choices.
Why it's a landmark: Showed a randomized bundle of housing-search assistance, not vouchers alone, sharply raised low-income families' moves to opportunity neighborhoods.
2024
Police Force Size and Civilian Race
Studies how changes in police force size relate to homicides and arrests across places, with estimated effects that differ by civilian race.
Why it's a landmark: Used federal hiring-grant variation to show added police reduce homicides while disproportionately raising low-level arrests of Black civilians.
2022
The Microgeography of Housing Supply
Develops a neighborhood-level framework to examine how housing-supply elasticities vary within and across metropolitan areas.
Why it's a landmark: Estimated housing-supply elasticities at neighborhood scale, showing within-metro variation dominates and that renovations and teardowns, not just new construction, drive responses.
2024
▾Labor Economics10 papers
Landmark Papers
The Impact of the Mariel Boatlift on the Miami Labor Market
Uses the Mariel Boatlift as a large, sudden immigration shock to study how a rise in labor supply affected wages and employment in Miami.
Why it's a landmark: Used the Mariel supply shock as a natural experiment, finding a sudden low-skill influx barely affected native wages or employment.
1990
Minimum Wages and Employment: A Case Study of the Fast-Food Industry in New Jersey and Pennsylvania
Uses New Jersey's minimum-wage increase as a natural experiment to study its effect on fast-food employment relative to neighboring Pennsylvania.
Why it's a landmark: Used a cross-border difference-in-differences on New Jersey's minimum-wage rise to find no employment loss, challenging competitive labor-demand predictions.
1994
Using Geographic Variation in College Proximity to Estimate the Return to Schooling
College proximity as IV — returns to education are higher for those induced to attend by proximity.
Why it's a landmark: Instrumented schooling with college proximity, estimating returns to education that exceeded OLS and reframed the ability-bias debate.
1995
Orchestrating Impartiality: The Impact of Blind Auditions on Female Musicians
Uses the adoption of blind orchestra auditions to study how screening procedures affect the advancement of women.
Why it's a landmark: Used the adoption of blind orchestra auditions as a natural experiment to identify sex bias in hiring evaluations.
2000
The Skill Content of Recent Technological Change
Routine-biased technological change — technology substitutes for routine tasks and complements cognitive ones.
Why it's a landmark: Introduced the routine-task framework, showing computers substitute for routine tasks and complement abstract ones, driving job polarization.
2003
Does Your High School Matter? Measuring Teacher Value-Added
Teacher quality has large long-run effects on earnings, college attendance, and teen birth rates.
Why it's a landmark: Validated teacher value-added against quasi-random turnover and linked it to students' long-run earnings and life outcomes.
2014
A Grand Gender Convergence: Its Last Chapter
The remaining gender pay gap is driven by the value of workplace flexibility, not discrimination alone.
Why it's a landmark: Argued the residual gender pay gap stems from nonlinear pay for long, inflexible hours rather than discrimination or human capital alone.
2014
Child Penalties Across Countries: Evidence and Explanations
Motherhood drives most of the gender earnings gap — child penalties are large and persistent across countries.
Why it's a landmark: Documented that earnings gaps open sharply at first childbirth, showing cross-country child penalties track gender norms more than policy.
2019
The Evolution of Work from Home
Uses survey data to examine why remote work persisted after the pandemic and how it varies across worker and job characteristics.
Why it's a landmark: Built real-time survey measures of remote work, documenting its persistent post-pandemic level and its estimated productivity and amenity value.
2023
The Economic Impacts of COVID-19: Evidence from a New Public Database Built Using Private Sector Data
Uses real-time private-sector data to examine how spending, business revenue, and low-wage employment responded to the COVID-19 shock across places.
Why it's a landmark: Built a public real-time private-data tracker of spending, employment, and revenue to measure the pandemic's granular economic impact.
2024
▾Behavioral & Computational7 papers
Landmark Papers
Prospect Theory: An Analysis of Decision under Risk
Foundations for how attention, loss aversion, and salience shape economic decisions.
Why it's a landmark: Introduced reference-dependent utility with loss aversion and probability weighting, supplanting expected-utility theory as the descriptive model of risky choice.
1979
Anomalies: The Endowment Effect, Loss Aversion, and Status Quo Bias
People demand more to give up an object than they would pay to acquire it — loss aversion shapes market behavior.
Why it's a landmark: Consolidated experimental evidence that ownership raises valuations, documenting the endowment effect and status-quo bias as consequences of loss aversion.
1991
Allocative Efficiency of Markets with Zero-Intelligence Traders
Institutions generate equilibrium outcomes even with minimally rational agents.
Why it's a landmark: Showed randomly bidding 'zero-intelligence' traders achieve near-efficient outcomes, attributing market efficiency to institutions rather than trader rationality.
1993
A Dynamic Model of an Individual's Income and Wealth
Heterogeneous agent model — precautionary savings and incomplete markets.
Why it's a landmark: Introduced the heterogeneous-agent general-equilibrium model with uninsurable idiosyncratic risk and borrowing constraints, founding the Bewley-Aiyagari framework.
1994
Attention Discrimination
Selective attention in screening — how limited attention amplifies discrimination.
Why it's a landmark: Used a field experiment to show that employers' costly information acquisition generates statistical discrimination through selective attention to applicants.
2016
Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence
Uses a randomized experiment to examine how access to a generative-AI writing tool affects task time, output quality, and within-worker inequality.
Why it's a landmark: Ran an early randomized trial showing a generative-AI writing tool raised productivity while compressing quality dispersion across workers.
2023
Large Language Models: An Applied Econometric Framework
Develops an econometric framework for using large language models in prediction and measurement tasks, examining issues such as training leakage and validation.
Why it's a landmark: Formalized when LLM outputs yield valid inference, requiring no training-data leakage for prediction and a validation-sample correction for estimation.
2026
▾Research Replicability & Transparency6 papers
Landmark Papers
Why Most Published Research Findings Are False
Foundational paper on false discovery rates — low power, researcher flexibility, and bias inflate Type I errors.
Why it's a landmark: Formalized how low prior odds, bias, and multiple testing make most published positive findings likely false.
2005
Estimating the Reproducibility of Psychological Science
Only 36% of psychology findings replicated — landmark call for reproducibility standards across sciences.
Why it's a landmark: Ran the first large-scale coordinated replication effort, finding under half of psychology studies replicated at reduced effect sizes.
2015
Star Wars: The Empirics Strike Back
Systematic evidence of p-hacking in economics — test statistics bunch just below conventional significance thresholds.
Why it's a landmark: Documented a two-humped distribution of test statistics around significance thresholds in economics journals, evidencing p-hacking and selective reporting.
2016
Evaluating Replicability of Laboratory Experiments in Economics
61% of economics lab experiments replicated — effect sizes in replications are about half of originals.
Why it's a landmark: Conducted coordinated high-powered replications of experimental economics studies, finding a majority replicated with somewhat attenuated effects.
2016
Identification of and Correction for Publication Bias
Structural model of publication bias — provides a method to correct estimates for selective reporting.
Why it's a landmark: Developed a method to estimate the selection function governing publication and reweight literatures to correct for the resulting bias.
2019
Methods Matter: p-Hacking and Publication Bias in Causal Inference
Publication bias is larger for IV and DiD designs than OLS — identification strategy affects selective reporting.
Why it's a landmark: Compared identification strategies across thousands of tests, showing IV and DiD exhibit more inflation from p-hacking and publication bias.
2020
▾Health Economics7 papers
Landmark Papers
Health Insurance Coverage and Medical Expenditures
RDD around Medicare eligibility at age 65 — insurance sharply reduces out-of-pocket spending.
Why it's a landmark: Used the Medicare-at-65 age discontinuity to identify sharp jumps in insurance coverage and healthcare utilization.
2008
Does Medicare Save Lives?
Near-elderly mortality — Medicare eligibility reduces mortality for low-income groups.
Why it's a landmark: Used the age-65 Medicare eligibility discontinuity among emergency admissions to identify a reduction in patient mortality.
2009
Menu Labeling and Consumer Choice
DiD around calorie labeling policy — consumer decisions respond to visible informational cues.
Why it's a landmark: Used a Starbucks natural experiment to show mandatory calorie posting modestly reduced calories purchased, concentrated among high-calorie consumers.
2011
The Oregon Health Insurance Experiment
Medicaid expansion by lottery — insurance affects access, utilization, and financial security.
Why it's a landmark: Exploited a Medicaid lottery for the first randomized evidence that coverage raised utilization and financial security but not measured physical health.
2012
Behavioral Hazard in Health Insurance
Patients underuse beneficial care — behavioral frictions distort health decisions beyond moral hazard.
Why it's a landmark: Introduced 'behavioral hazard,' showing cost-sharing can cut high-value care when patients underuse it, complicating the standard moral-hazard logic.
2015
Place-Based Drivers of Mortality: Evidence from Migration
Uses Medicare movers to examine how much geographic variation in elderly mortality reflects current place versus person-specific health.
Why it's a landmark: Used Medicare movers to separate place from person, showing location causally affects mortality largely through healthcare use.
2021
Social Capital I: Measurement and Associations with Economic Mobility
Uses large-scale social-network data to construct measures of social capital and examine their association with upward economic mobility across areas.
Why it's a landmark: Used Facebook-scale networks to measure social capital, showing cross-class friendship ('economic connectedness') strongly predicts upward mobility.
2022
▾Causal Inference & Methods12 papers
Landmark Papers
Let's Take the Con Out of Econometrics
Classic critique of specification searching and data mining — a foundational call for credible empirical practice.
Why it's a landmark: Argued regression inference hinges on fragile priors, introducing extreme-bounds sensitivity analysis and catalyzing the drive toward credible identification.
1983
Identification and Estimation of Local Average Treatment Effects
The LATE framework — IV identifies effects for compliers, not the full population.
Why it's a landmark: Introduced the LATE framework, showing instrumental variables identify treatment effects only for compliers under a monotonicity assumption.
1994
Identification of Causal Effects Using Instrumental Variables
Foundational paper on the potential outcomes framework for IV.
Why it's a landmark: Recast IV within the Rubin potential-outcomes model, formalizing the exclusion, monotonicity, and independence assumptions underlying causal instrument estimates.
1996
Are Emily and Greg More Employable than Lakisha and Jamal?
Résumé audit experiment — institutional screening generates persistent labor market inequities.
Why it's a landmark: Introduced the resume-audit correspondence experiment, documenting large callback gaps from randomly assigned racially distinctive names.
2004
The Credibility Revolution in Empirical Economics
How better identification strategies — RCTs, IV, RDD, DiD — transformed empirical economics into a credible science.
Why it's a landmark: Codified the design-based 'credibility revolution,' arguing research-design transparency, not structural modeling, secured empirical economics' reliability.
2010
Double/Debiased Machine Learning for Treatment and Structural Parameters
DML — uses cross-fitting and Neyman orthogonality to combine ML with causal inference.
Why it's a landmark: Developed Neyman-orthogonal, cross-fitted estimators that let machine-learning nuisance estimates deliver valid inference on low-dimensional causal parameters.
2018
Difference-in-Differences with Multiple Time Periods
Group-time ATT framework — robust to TWFE bias under staggered adoption.
Why it's a landmark: Introduced group-time average treatment effects with a clean control group, providing robust DiD estimators under staggered adoption and heterogeneity.
2021
Two-Way Fixed Effects Estimators with Heterogeneous Treatment Effects
Studies how two-way fixed-effects regressions weight group-period treatment effects and proposes an alternative estimator under heterogeneity.
Why it's a landmark: Showed two-way fixed-effects estimates are contaminated 'negative-weight' averages under heterogeneous effects, and proposed a robust alternative estimator.
2020
Difference-in-Differences with Variation in Treatment Timing
Decomposes the two-way fixed-effects estimator into a weighted average of all two-group/two-period comparisons to examine timing-driven bias.
Why it's a landmark: Decomposed staggered DiD into all two-by-two comparisons, exposing the bias from using already-treated units as controls.
2021
Estimating Dynamic Treatment Effects in Event Studies with Heterogeneous Treatment Effects
Studies how event-study lead and lag coefficients can be contaminated under heterogeneous timing and proposes an interaction-weighted estimator.
Why it's a landmark: Showed event-study coefficients mix effects across cohorts under heterogeneity, and proposed an interaction-weighted estimator for clean dynamic estimates.
2021
A More Credible Approach to Parallel Trends
Develops sensitivity-analysis tools for difference-in-differences and event studies that relax exact parallel trends by bounding post-treatment violations.
Why it's a landmark: Replaced binary parallel-trends tests with a bounded sensitivity analysis, deriving robust confidence sets under formal restrictions on trend violations.
2023
Revisiting Event-Study Designs: Robust and Efficient Estimation
Develops an imputation-based estimator for staggered event-study designs and examines its efficiency and robustness under treatment-effect heterogeneity.
Why it's a landmark: Developed an imputation estimator that recovers untreated potential outcomes from controls, yielding efficient, robust event-study estimates under heterogeneity.
2024
