Update on a regression discontinuity dispute: some asynchronous collaboration

Statistical Modeling, Causal Inference, and Social Science 2026-08-19

Anjali Thomas writes:

I am writing to share a paper which is a re-examination of my 2018 AJPS article entitled “Targeting Ordinary Voters or Political Elites”  which was previously discussed on this blog here.

The paper, written in the spirit of Gelman (2022), conducts a thorough re-analysis of my earlier work in light of recent developments in regression discontinuity designs (RDD). It also directly addresses specific critiques of the article raised both in subsequent academic literature and in previous comments on this blog.

A full response to each critique raised on this blog appears in Section A.3 on page 48 of the paper. Among other things, the blog critiqued the use of the global fourth order polynomial, and commenters highlighted that it appeared to be picking up noise in the data rather than a true relationship. While the original article did present results in the SI showing robustness to a local-linear specification with alternative bandwidths, the currentpaper significantly extends these checks and presents new results showing:
  • Polynomial & Bandwidth Robustness: Shows stability across lower-order polynomials, alternative bandwidths, kernel choices, and a donut-hole approach.

  • Noise Reduction via Aggregation: Aggregates data to the level of the running variable to reduce noise, presenting new scatter plots where the visual discontinuity persists at this level of aggregation.

  • Spatial Structure: Demonstrates that reported balance/placebo issues in recent re-analyses stem from ignoring within-constituency clustering and predictive controls.

A link to the paper is here (https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7272598) and the abstract is below:

This note reexamines Thomas (2018), which advances and tests the ‘elite cooperation logic’ whereby national politicians target resources along partisan lines to win over the cooperation of co-partisan state legislators in implementing development projects. Consistent with this logic, a close election regression discontinuity design (RDD) uses project-level data to show that national legislators in North India allocate systematically higher public works expenditures to constituencies of co-partisan state legislators in the period after a state election. Responding to critiques in subsequent literature of the RDD approach used, this note shows that both the evidence of covariate imbalance reported in Bicalho et al. (2026) and the high proportion of significant placebo estimates reported in Albada (2025) are artefacts of ignoring within-constituency clustering and omitting predictive controls. Meanwhile, this note presents re-analyses confirming that the core findings in Thomas (2018) are robust to covariate inclusion using either constituency-level clustering or constituency-level aggregation. Acknowledging the problems related to global fourth order polynomials (Gelman and Imbens, 2019; Albada, 2025), the results in Thomas (2018) are also shown to be robust to lower-order polynomials, alternative bandwidths, alternative kernel choices, the donut hole approach, and an inference procedure that adjusts for worst-case bias (Stommes et al., 2023). The note re-establishes the credibility of the substantive findings in Thomas (2018) and highlights the importance of accounting for spatial clustering in both estimation and diagnostic checks in RDDs.

Anjali took my Bayesian statistics class back in 2006! It’s great to see what former students are doing, and I love seeing this sort of asynchronous collaboration.

Regarding the regression discontinuity issues, I do not think it makes sense to fit unrelated curves on the two sides of the cutoff. So I prefer the versions that fit a single curve plus discontinuity, not two curves.

The other thing I always recommend (see this recent paper with Imbens is that regressions include not just the forcing variable but also other pre-treatment predictors. This should help with both bias and efficiency.

All the discussion of bandwidth, functional forms, etc., can be wasted if the fitted model makes no sense or if it’s a distraction from including additional pre-treatment predictors.

In any case, it’s great to see this sort of open exploration.