Skip to content

fix(synthetic_control): put the treated unit in every in-space placebo's donor pool - #841

Open
ilgrad wants to merge 1 commit into
igerber:mainfrom
ilgrad:fix/in-space-placebo-donor-pool
Open

ilgrad wants to merge 1 commit into
igerber:mainfrom
ilgrad:fix/in-space-placebo-donor-pool

Conversation

@ilgrad

@ilgrad ilgrad commented Oct 5, 2026

Copy link
Copy Markdown

Fixes #837.

Problem

SyntheticControlResults.in_space_placebo() fits each placebo against the other J − 1 donors and never puts the real treated unit in a placebo's donor pool, while the treated unit's own synthetic control uses all J donors. Unit j's RMSPE ratio then depends on which unit carries the treatment label, so the J + 1 ratios are not one fixed set that random assignment permutes, and the permutation p-value is not exact. On the #837 panel (10 units, no effect) 2 of the 10 assignments give p = 0.1 and 3 give p ≤ 0.2. ADH 2010 build every placebo from all the other units (§3.4, journal p. 501): "In each iteration we reassign in our data the tobacco control intervention to one of the 38 control states, shifting California to the donor pool." Firpo & Possebom (2018, fn 10) keep the treated unit in the placebo pools too. test_sharp_null() and confidence_set() re-rank the same placebo gaps, so they inherit the construction.

Reproducer

main 6874b5d, Python 3.14.7, numpy 2.5.3, pandas 3.0.6, scipy 1.18.1, pure-Python backend (DIFF_DIFF_BACKEND=python). The script is the one in #837, with two lines added at the end to print the per-unit p-values.

Script
import warnings

import numpy as np
import pandas as pd
from diff_diff import SyntheticControl

N_PRE, N_POST, UNITS = 24, 6, 10


def panel(seed):
    rng = np.random.default_rng(seed)
    periods = N_PRE + N_POST
    factors = np.empty((periods, 3))
    errors = np.empty((UNITS, periods))
    factors[0] = rng.normal(size=3) / np.sqrt(0.75)
    errors[:, 0] = rng.normal(size=UNITS) / np.sqrt(0.75)
    for t in range(1, periods):
        factors[t] = 0.5 * factors[t - 1] + rng.normal(size=3)
        errors[:, t] = 0.5 * errors[:, t - 1] + rng.normal(size=UNITS)
    loadings = rng.uniform(0.0, 1.0, (UNITS, 3))
    return loadings @ factors.T + 0.5 * errors + rng.normal(size=(UNITS, 1))


def fit(y, treated):
    d = np.zeros(y.shape, dtype=int)
    d[treated, N_PRE:] = 1
    frame = pd.DataFrame({"unit": np.repeat(np.arange(UNITS), y.shape[1]),
                          "time": np.tile(np.arange(y.shape[1]), UNITS), "y": y.ravel(), "d": d.ravel()})
    # Uniform V and a generous Frank-Wolfe budget, so every fit converges and only the
    # placebos' donor pools differ between the two constructions.
    return SyntheticControl(v_method="custom", custom_v=np.ones(N_PRE), standardize="none",
                            inner_max_iter=200_000, seed=0).fit(
        frame, outcome="y", treatment="d", unit="unit", time="time", pre_period_outcomes="all")


y = panel(0)
with warnings.catch_warnings():
    warnings.simplefilter("ignore")  # poor-pre-fit notices on some units; not the point here
    fits = [fit(y, unit) for unit in range(UNITS)]
    for f in fits:
        f.in_space_placebo()
assert all(f.n_failed == 0 for f in fits)

p_in_space = np.array([f.placebo_p_value for f in fits])
# ADH's construction from diff-diff's own fits: fits[u] is unit u's synthetic control from all the others.
ratio = np.array([f.rmspe_ratio for f in fits])
p_all_others = np.array([np.mean(ratio >= ratio[unit]) for unit in range(UNITS)])
print("in_space_placebo():         ", np.round(np.sort(p_in_space), 2))
print("each unit from all others:  ", np.round(np.sort(p_all_others), 2))
for k in range(1, UNITS + 1):
    count = int(np.sum(p_in_space <= k / UNITS + 1e-12))
    if count > k:
        print(f"in_space_placebo(): {count} of the {UNITS} assignments give p <= {k}/{UNITS}")
print("per-unit in_space_placebo p:", p_in_space.tolist())
print("per-unit all-others p:      ", p_all_others.tolist())

On main:

in_space_placebo():          [0.1 0.1 0.2 0.5 0.5 0.6 0.7 0.8 0.9 1. ]
each unit from all others:   [0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1. ]
in_space_placebo(): 2 of the 10 assignments give p <= 1/10
in_space_placebo(): 3 of the 10 assignments give p <= 2/10
per-unit in_space_placebo p: [0.2, 0.6, 0.1, 0.5, 0.5, 1.0, 0.8, 0.9, 0.7, 0.1]
per-unit all-others p:       [0.3, 0.7, 0.1, 0.5, 0.4, 1.0, 0.8, 0.9, 0.6, 0.2]

On this branch:

in_space_placebo():          [0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1. ]
each unit from all others:   [0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1. ]
per-unit in_space_placebo p: [0.3, 0.7, 0.1, 0.5, 0.4, 1.0, 0.8, 0.9, 0.6, 0.2]
per-unit all-others p:       [0.3, 0.7, 0.1, 0.5, 0.4, 1.0, 0.8, 0.9, 0.6, 0.2]

Each placebo refit is now the fit that treats that unit, so in_space_placebo() reproduces the all-others p-values exactly. Five of the ten assignments move by one rank (three up, two down).

Cause

diff_diff/synthetic_control_results.py:1030 on main, in in_space_placebo(): pool = [d for d in donors if d != j], where donors = snap.donor_ids never contains the treated unit. The docstring and the REGISTRY Note gave two reasons: the treated unit's post period is treatment-contaminated, and the construction matches SCtools::generate.placebos. The first holds only under the alternative. Under the null the p-value is calibrated for, the treated unit's outcome is an untreated outcome like any other. Exactness needs each unit's statistic to be the same function of the data whichever unit is labelled treated, and leaving the label's unit out of every other pool breaks exactly that.

Fix

  • in_space_placebo(): placebo j's pool is every unit except j, the treated unit included, in the outcome pivot's sorted unit order. A unit's refit therefore sees the same pool in the same order whichever unit is treated, and with the default donor pool it is the same computation as a fresh fit() that treats that unit (the regression test checks the ratios agree to rtol=1e-12).
  • _placebo_fit_unit() (private) also returns the donor weights. leave_one_out() and in_time_placebo() ignore them.
  • test_sharp_null() / confidence_set(): under H_0^f the treated unit's untreated outcome is Y_1 − f, so placebo j's synthetic control moves by −w_j1·f, where w_j1 is its weight on the treated unit (Firpo & Possebom fn 10: "under sharp null hypotheses, the treatment effect is known"). in_space_placebo() keeps each placebo's w_j1; _sharp_null_pvalue() and _invert_sharp_null() take them as a keyword-only treated_weights and use the placebo residual g_j − (1 − w_j1)·f. The breakpoint quadratics get the matching (1 − w_j1)²·S2 curvature. With w_j1 = 0 this is the existing Eq 12 shift, so the existing helper tests are unchanged. The weights depend on pre-period data only, so this needs no refits; a brute-force test refits every placebo on the null-imputed panel and compares.
  • Messages that described the old pool: the J == 2 warning (each placebo is now fit against 2 units, with a weight solve and a V search), and the J < 2 reasons in the warning, summary(), _require_placebo_reference() and DiagnosticReport (with one donor, the treated unit and the donor are each other's synthetic control; this case still returns NaN).
  • Docs: the in_space_placebo() / test_sharp_null() / helper docstrings; REGISTRY §SyntheticControl (the donor-pool Note now cites ADH §3.4 p. 501, a new **Deviation from R (SCtools::generate.placebos):** entry, the test-inversion paragraph, the reference-implementation line, the checklist, and the validation Note); guides/llms-full.txt; one sentence in the Firpo-Possebom paper review; a changelog fragment.

Compatibility

  • Results. placebo_p_value, the placebo rows of get_placebo_df(), test_sharp_null() and confidence_set() change on every fit where some placebo's synthetic control puts weight on the treated unit. att, the donor weights, the treated unit's rmspe_ratio, leave_one_out() and in_time_placebo() do not change.
  • Cost. Each placebo's weight solve has J donors instead of J − 1.
  • J == 2. Each placebo is fit against 2 units, with a weight solve and a V search, instead of a forced w = [1]. With fewer than 2 donors the p-value is still NaN.
  • API. No public signature changes. Private helpers only: _placebo_fit_unit() also returns the weights, and _sharp_null_pvalue() / _invert_sharp_null() take a keyword-only treated_weights whose default None keeps the old shift.
  • R parity. The placebo pools no longer match SCtools::generate.placebos; REGISTRY records it as a **Deviation from R** entry.

Methodology references (required if estimator / math changes)

  • Method name(s): Abadie, Diamond & Hainmueller (2010) in-space placebo permutation inference; Firpo & Possebom (2018) sharp-null test inversion.
  • Paper / source link(s): https://doi.org/10.1198/jasa.2009.ap08746 (§3.4, journal p. 501), https://doi.org/10.1515/jci-2016-0026 (fn 10, Eqs 11–14)
  • Any intentional deviations from the source (and why): none from the papers; this removes one. The deviation is now from SCtools::generate.placebos, which drops the treated unit from the placebo pools; its reference set changes with the treated unit, so its p-value is not exact. Documented as a REGISTRY **Deviation from R** entry.

Validation

  • Tests added/updated, in tests/test_methodology_synthetic_control.py:
    • test_in_space_placebo_puts_real_treated_in_every_donor_pool: each placebo's pool is every unit but itself.
    • test_in_space_placebo_is_exact_permutation: on a 10-unit no-effect panel the ten assignments give p = 0.1, 0.2, ..., 1.0 once each, and each placebo's ratio equals that unit's own fit() to rtol=1e-12.
    • test_sharp_null_pvalue_imputes_treated_in_placebo_pools and test_invert_with_treated_weights_matches_direct_pvalues: the (1 − w_j1) shift in the p-value and in the inversion's breakpoints, against direct p-values.
    • test_test_sharp_null_matches_refits_on_null_imputed_panel: brute force. It refits every placebo on the panel whose treated post outcomes are replaced by Y_1 − f, and compares with test_sharp_null().
    • Updated: test_leave_one_out_uniform_shift_surfaced_by_delta_not_range (its stub of _placebo_fit_unit returns the weights too) and test_cv_in_space_placebo_excludes_donor_flat_refits (the treated unit now takes d1–d3's values in the validation window, so d0's placebo pool, which now holds the treated unit, stays donor-indistinguishable as the test intends).
  • Fail before, pass after. On main 6874b5d the two behavioural tests fail on their assertions: the pools lack treated, and the p-values are [0.2, 0.2, 0.2, 0.5, 0.5, 0.6, 0.7, 0.8, 1, 1]. The three sharp-null tests and the leave-one-out stub fail on the new private interface. The cv fixture test passes on both. All seven pass on this branch.
  • Mutations. Each of six fails at least one test: the p-value ignoring w_j1 (3 tests fail), the inversion ignoring it (1), the p-value's curvature without (1 − w_j1)² (4), the breakpoints' curvature without it (1), the old pool (40), the treated unit first in the pool's order (1).
  • Simulation evidence: three AR(1) factors with U(0, 1) loadings, AR(1) noise, unit levels, 20 pre and 5 post periods, uniform custom V.
    • Size, 300 no-effect panels of 10 units, each unit treated in turn, so the treated draw is integrated out exactly:

      p ≤ 0.1 p ≤ 0.2
      main 11.70% (MC SE 0.25 points); 2 of 10 units on 57 panels, none on 6 21.73% (SE 0.31); 3 or 4 of 10 on 71 panels, 1 on 20
      this branch 10% on every panel 20% on every panel
    • Placebo refits against fresh fits: on 40 panels, and on every panel whose drawn treated unit the branch rejects at 10% (342 panels of 10 units, 182 of 20), each placebo's ratio equals that unit's own fit() exactly (relative difference 0.0), and the stored p-value equals the p-value computed from the fresh fits.

    • The default nested V search: a placebo refit calls the same _outer_solve_V with the same seed and n_starts as a fresh fit, over the same units in the same order. On a panel of 5 units, 8 pre and 2 post periods (seed=3), each placebo's ratio equals that unit's own fit() exactly (relative difference 0.0), no placebo failed, and the p-value, 0.8, equals the share of the five fresh fits' ratios at or above the treated unit's.

  • Commands: Python 3.14.7, numpy 2.5.3, pandas 3.0.6, scipy 1.18.1, DIFF_DIFF_BACKEND=python.
    • pytest -n 2 on test_methodology_synthetic_control.py, test_t25_synthetic_control_policy_drift.py, test_diagnostic_report.py, test_business_report.py, test_practitioner.py, test_changelog_fragments.py, test_guides.py and test_doc_deps_integrity.py: 1158 passed, 1 skipped.
    • The slow test_confidence_set_coverage_simulation: passed. test_naming_guard.py, test_docs_ia.py, test_dml_carousel_claims.py and the cv fixture test: 59 passed.
    • ruff check diff_diff tests (0.16.7) and black --check diff_diff tests (26.5.1): clean. mypy diff_diff (2.3.1, with numpy 2.4.5, pandas 3.0.3, scipy 1.17.1, as in lint.yml): no issues.
    • Not run: the rest of the suite, the Rust backend, the notebooks.

Security / privacy

  • Confirm no secrets/PII in this PR: confirmed. Code, tests and docs only; the simulations use synthetic data.

Changelog

…o's donor pool

in_space_placebo() fit each placebo against the other J - 1 donors and never
put the real treated unit in a placebo's pool, while the treated unit's own
synthetic control uses all J donors. Unit j's RMSPE ratio then depended on
which unit carried the treatment label, so the J + 1 ratios were not one set
that random assignment permutes, and the permutation p-value was not exact.
Abadie, Diamond & Hainmueller (2010, section 3.4) shift the treated unit into
the donor pool. Each placebo now fits against every other unit, and
test_sharp_null() / confidence_set() use the treated unit's null-imputed
outcome Y_1 - f inside each placebo's synthetic control (Firpo & Possebom
2018, fn 10), weighted by the placebo's weight on it.

Fixes igerber#837

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: in_space_placebo() leaves the treated unit out of every placebo's donor pool, so the placebo test is not exact

1 participant