Repository navigation
Conversation
…o's donor pool in_space_placebo() fit each placebo against the other J - 1 donors and never put the real treated unit in a placebo's pool, while the treated unit's own synthetic control uses all J donors. Unit j's RMSPE ratio then depended on which unit carried the treatment label, so the J + 1 ratios were not one set that random assignment permutes, and the permutation p-value was not exact. Abadie, Diamond & Hainmueller (2010, section 3.4) shift the treated unit into the donor pool. Each placebo now fits against every other unit, and test_sharp_null() / confidence_set() use the treated unit's null-imputed outcome Y_1 - f inside each placebo's synthetic control (Firpo & Possebom 2018, fn 10), weighted by the placebo's weight on it. Fixes igerber#837
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #837.
Problem
SyntheticControlResults.in_space_placebo()fits each placebo against the otherJ − 1donors and never puts the real treated unit in a placebo's donor pool, while the treated unit's own synthetic control uses allJdonors. Unitj's RMSPE ratio then depends on which unit carries the treatment label, so theJ + 1ratios are not one fixed set that random assignment permutes, and the permutation p-value is not exact. On the #837 panel (10 units, no effect) 2 of the 10 assignments givep = 0.1and 3 givep ≤ 0.2. ADH 2010 build every placebo from all the other units (§3.4, journal p. 501): "In each iteration we reassign in our data the tobacco control intervention to one of the 38 control states, shifting California to the donor pool." Firpo & Possebom (2018, fn 10) keep the treated unit in the placebo pools too.test_sharp_null()andconfidence_set()re-rank the same placebo gaps, so they inherit the construction.Reproducer
main
6874b5d, Python 3.14.7, numpy 2.5.3, pandas 3.0.6, scipy 1.18.1, pure-Python backend (DIFF_DIFF_BACKEND=python). The script is the one in #837, with two lines added at the end to print the per-unit p-values.Script
On main:
On this branch:
Each placebo refit is now the fit that treats that unit, so
in_space_placebo()reproduces the all-others p-values exactly. Five of the ten assignments move by one rank (three up, two down).Cause
diff_diff/synthetic_control_results.py:1030on main, inin_space_placebo():pool = [d for d in donors if d != j], wheredonors = snap.donor_idsnever contains the treated unit. The docstring and the REGISTRY Note gave two reasons: the treated unit's post period is treatment-contaminated, and the construction matchesSCtools::generate.placebos. The first holds only under the alternative. Under the null the p-value is calibrated for, the treated unit's outcome is an untreated outcome like any other. Exactness needs each unit's statistic to be the same function of the data whichever unit is labelled treated, and leaving the label's unit out of every other pool breaks exactly that.Fix
in_space_placebo(): placeboj's pool is every unit exceptj, the treated unit included, in the outcome pivot's sorted unit order. A unit's refit therefore sees the same pool in the same order whichever unit is treated, and with the default donor pool it is the same computation as a freshfit()that treats that unit (the regression test checks the ratios agree tortol=1e-12)._placebo_fit_unit()(private) also returns the donor weights.leave_one_out()andin_time_placebo()ignore them.test_sharp_null()/confidence_set(): underH_0^fthe treated unit's untreated outcome isY_1 − f, so placeboj's synthetic control moves by−w_j1·f, wherew_j1is its weight on the treated unit (Firpo & Possebom fn 10: "under sharp null hypotheses, the treatment effect is known").in_space_placebo()keeps each placebo'sw_j1;_sharp_null_pvalue()and_invert_sharp_null()take them as a keyword-onlytreated_weightsand use the placebo residualg_j − (1 − w_j1)·f. The breakpoint quadratics get the matching(1 − w_j1)²·S2curvature. Withw_j1 = 0this is the existing Eq 12 shift, so the existing helper tests are unchanged. The weights depend on pre-period data only, so this needs no refits; a brute-force test refits every placebo on the null-imputed panel and compares.J == 2warning (each placebo is now fit against 2 units, with a weight solve and aVsearch), and theJ < 2reasons in the warning,summary(),_require_placebo_reference()and DiagnosticReport (with one donor, the treated unit and the donor are each other's synthetic control; this case still returns NaN).in_space_placebo()/test_sharp_null()/ helper docstrings; REGISTRY §SyntheticControl (the donor-pool Note now cites ADH §3.4 p. 501, a new**Deviation from R (SCtools::generate.placebos):**entry, the test-inversion paragraph, the reference-implementation line, the checklist, and the validation Note);guides/llms-full.txt; one sentence in the Firpo-Possebom paper review; a changelog fragment.Compatibility
placebo_p_value, the placebo rows ofget_placebo_df(),test_sharp_null()andconfidence_set()change on every fit where some placebo's synthetic control puts weight on the treated unit.att, the donor weights, the treated unit'srmspe_ratio,leave_one_out()andin_time_placebo()do not change.Jdonors instead ofJ − 1.J == 2. Each placebo is fit against 2 units, with a weight solve and aVsearch, instead of a forcedw = [1]. With fewer than 2 donors the p-value is still NaN._placebo_fit_unit()also returns the weights, and_sharp_null_pvalue()/_invert_sharp_null()take a keyword-onlytreated_weightswhose defaultNonekeeps the old shift.SCtools::generate.placebos; REGISTRY records it as a**Deviation from R**entry.Methodology references (required if estimator / math changes)
SCtools::generate.placebos, which drops the treated unit from the placebo pools; its reference set changes with the treated unit, so its p-value is not exact. Documented as a REGISTRY**Deviation from R**entry.Validation
tests/test_methodology_synthetic_control.py:test_in_space_placebo_puts_real_treated_in_every_donor_pool: each placebo's pool is every unit but itself.test_in_space_placebo_is_exact_permutation: on a 10-unit no-effect panel the ten assignments givep = 0.1, 0.2, ..., 1.0once each, and each placebo's ratio equals that unit's ownfit()tortol=1e-12.test_sharp_null_pvalue_imputes_treated_in_placebo_poolsandtest_invert_with_treated_weights_matches_direct_pvalues: the(1 − w_j1)shift in the p-value and in the inversion's breakpoints, against direct p-values.test_test_sharp_null_matches_refits_on_null_imputed_panel: brute force. It refits every placebo on the panel whose treated post outcomes are replaced byY_1 − f, and compares withtest_sharp_null().test_leave_one_out_uniform_shift_surfaced_by_delta_not_range(its stub of_placebo_fit_unitreturns the weights too) andtest_cv_in_space_placebo_excludes_donor_flat_refits(the treated unit now takes d1–d3's values in the validation window, so d0's placebo pool, which now holds the treated unit, stays donor-indistinguishable as the test intends).6874b5dthe two behavioural tests fail on their assertions: the pools lacktreated, and the p-values are[0.2, 0.2, 0.2, 0.5, 0.5, 0.6, 0.7, 0.8, 1, 1]. The three sharp-null tests and the leave-one-out stub fail on the new private interface. The cv fixture test passes on both. All seven pass on this branch.w_j1(3 tests fail), the inversion ignoring it (1), the p-value's curvature without(1 − w_j1)²(4), the breakpoints' curvature without it (1), the old pool (40), the treated unit first in the pool's order (1).V.Size, 300 no-effect panels of 10 units, each unit treated in turn, so the treated draw is integrated out exactly:
p ≤ 0.1p ≤ 0.2Placebo refits against fresh fits: on 40 panels, and on every panel whose drawn treated unit the branch rejects at 10% (342 panels of 10 units, 182 of 20), each placebo's ratio equals that unit's own
fit()exactly (relative difference 0.0), and the stored p-value equals the p-value computed from the fresh fits.The default nested
Vsearch: a placebo refit calls the same_outer_solve_Vwith the sameseedandn_startsas a fresh fit, over the same units in the same order. On a panel of 5 units, 8 pre and 2 post periods (seed=3), each placebo's ratio equals that unit's ownfit()exactly (relative difference 0.0), no placebo failed, and the p-value, 0.8, equals the share of the five fresh fits' ratios at or above the treated unit's.DIFF_DIFF_BACKEND=python.pytest -n 2ontest_methodology_synthetic_control.py,test_t25_synthetic_control_policy_drift.py,test_diagnostic_report.py,test_business_report.py,test_practitioner.py,test_changelog_fragments.py,test_guides.pyandtest_doc_deps_integrity.py: 1158 passed, 1 skipped.test_confidence_set_coverage_simulation: passed.test_naming_guard.py,test_docs_ia.py,test_dml_carousel_claims.pyand the cv fixture test: 59 passed.ruff check diff_diff tests(0.16.7) andblack --check diff_diff tests(26.5.1): clean.mypy diff_diff(2.3.1, with numpy 2.4.5, pandas 3.0.3, scipy 1.17.1, as inlint.yml): no issues.Security / privacy
Changelog
changelog.d/20261005-in-space-placebo-donor-pool.md,### Fixed. It states thatplacebo_p_valueand the placebo rows ofget_placebo_df()change on every fit where a placebo puts weight on the treated unit, gives the direction and size on the [Bug]: in_space_placebo() leaves the treated unit out of every placebo's donor pool, so the placebo test is not exact #837 panel, and notes thattest_sharp_null()/confidence_set()change too.