Repository navigation
Conversation
…ve least squares Frank-Wolfe with a budget of inner_max_iter steps stopped short of the simplex least-squares optimum on ordinary factor-model panels, so the treated fit failed closed, placebos were dropped as failed, and the nested V search ranked V by where truncated solves stopped. The problem reduces exactly to one non-negative least squares, solved here by Lawson-Hanson's active set, which terminates finitely and, being diff-diff's own, returns the same vertex at a tie under every SciPy release. The CWZ conformal proxy and rank_control_units' synthetic_weight column solve the same problem and use the same solve. Fixes igerber#838
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #838.
Problem
SyntheticControlsolved its donor weights,min_W ||V^½ (X1 − X0 W)||²over the simplex, by Frank-Wolfe with a budget ofinner_max_iter=10000steps. Frank-Wolfe converges sublinearly when the optimum lies on a face of the simplex, the usual case for synthetic control weights, and on ordinary factor-model panels it stopped on that budget. The treated fit then warned "did not converge" and the diagnostics that need a converged fit were skipped (placebo_p_valueNaN); placebos whose refit hit the budget were dropped from the reference set; and the nestedVsearch rankedVonly among the solves that met their stopping rule. In #838's reproducer the placebo test failed closed or lost placebos on 40, 40 and 33 of 40 panels, and at the full defaults the nested search excluded 46%, 25% and 0.5% of its solves fromVranking.Fix
The problem reduces exactly to one non-negative least squares (the derivation is in #838 and in
_sc_weight_nnls's docstring): withZ = V^½ (X0 − X1 1'),u = argmin_{u ≥ 0} ||[Z / s; 1'] u − [0; 1]||andW = u / 1'u, for any scales > 0. An active-set method solves it in finitely many steps, so there is no budget to exhaust.utils._sc_weight_nnls(Y, zeta):_sc_weight_fw's uncentered problem (intercept=False),zeta² ||w||² + ||A w − b||² / Non the simplex, solved exactly.zeta > 0stacks√N ζ IunderZ.utils._nnls_lawson_hanson(M, d, max_iter): Lawson and Hanson's active set (1974, ch. 23), with their guard against an entering column that rounding leaves without positive weight. It is ~40 lines of NumPy rather than a call toscipy.optimize.nnls; see SciPy versions. Pastmax_iter(50 per donor) it raisesRuntimeErrorrather than return a truncated solve; no handler in these modules catches it, and no test or run below raised it._inner_solve_W, which everyVevaluation and every placebo, leave-one-out, in-time and sparse refit goes through, calls_sc_weight_nnlsand always reports converged.conformal._cwz_proxy_fit,zeta=0), andprep.rank_control_units's informationalsynthetic_weightcolumn (zeta = sqrt(lambda_reg / N), which stopped atmax_iter=1000). A panel with a missing pre-period outcome keeps the uniform column it had.inner_max_iter/inner_min_decreaseas the remedy, now name the outer search's knobs (optimizer_options['maxiter'],n_starts), or none where only the inner solve could fail.synthdidparity) is untouched.Choices
The four from my comment on #838:
Won a face of the simplex fits them exactly. The active set returns a vertex of that face (at mostk + 1donors); RSynth'sipoptends to a point inside it, and Frank-Wolfe stopped wherever its budget left it. REGISTRY records this as a**Note (deviation from R)**. I added no selection rule. I tried the minimum-normW(a 1e-10 relative ridge) on Tutorial 25:W(V)becomes continuous, the nested search then drivesVto(0.017, 2e-12, 0.983), a corner where one predictor's weight is negligible, and the ATT to 6.91. A TODO row records the question; Abadie and L'Hour (2021) is the published rule.inner_max_iter,inner_min_decreaseare accepted and validated but no longer used, and their docstrings say so. Deprecating them is a TODO row._inner_solve_Wto report a truncated solve (_truncate_inner_solves). Keeping or removing the paths is in the same TODO row.SciPy versions
At a tie, the vertex an active set returns depends on its pivoting, and
scipy.optimize.nnlshas been reimplemented several times. With it, Tutorial 25's 21 placebo RMSPE ratios took five different sets over SciPy 1.11 to 1.18, where main's Frank-Wolfe ratios agree to 2e-5, and under 1.15 the constant-effect confidence set was unbounded. With_nnls_lawson_hansonthe ratios agree to the sixth decimal under 1.11 to 1.18.Tutorial 25 under each release,
DIFF_DIFF_BACKEND=python, Python 3.11 (1.17.1 on 3.12, 1.18.1 on 3.14). A letter names one set of the 21 placebo RMSPE ratios, compared at six decimals; the brackets are the constant-effect set as the tutorial prints it.scipy.optimize.nnlsA* is A with one ratio 1.5e-5 away (1.119307 → 1.119322). A′ and B′ each move one placebo under SciPy 1.10; that comes from the outer
Vsearch, not the inner solve, on main as on this branch.The constant set's printed endpoints move by one grid step between environments on both branches (main: 4.16 in the first seven rows, 4.05 in the last two). That is older than this PR: the tutorial prints the first and last accepted points of the 200-point display grid that
_invert_sharp_nulllays between the exact endpoints, and at those two points a placebo ties the treated unit, so rounding decides whether they show as accepted. The exact endpoints ineffect_confidence_setdo not depend on that tie. I left it alone here.Behavioral changes
Estimates move by the former solver's gap. On [Bug]: SyntheticControl's Frank-Wolfe stops on its budget at the default inner_max_iter; in_space_placebo() then fails closed or drops placebos on 33-40 of 40 panels #838's panels the treated pre-RMSPE was a median 1e-4 above the optimum (relative); it now equals it to 7e-16. At a tie the move can be large, because Frank-Wolfe's stopping point and the vertex are different exact fits. Tutorial 25 is such a case, and its text and drift bands are updated:
The tutorial's prose still holds (every in-time placebo below 1 in absolute value, the leave-one-out bound, the conformal claims); the numbers quoted in it are updated.
Conformal intervals are narrower and still cover above nominal. Over 50 simulated panels (6 donors, 22 periods, constant effect 6; 300 period intervals at nominal 90%, inverted on a fine grid), coverage goes from 0.970 to 0.953 and the median width from 0.515 to 0.450. No grid point failed to converge on either branch.
One existing test now inverts on a fine grid:
test_conformal_confidence_intervals_recover_true_constant_effect. The default grid puts 100 points on±0.5·|estimate|, a step of about 0.06 there, against intervals about 0.2 wide, so each interval's endpoints sit up to a step inside the accepted set. On that grid this branch brackets the true effect in 4 of 6 periods (the test asks for 5); onbounds=(5, 7), n_grid=401it brackets 5 of 6, and main 6 of 6 on both grids. A TODO row records the grid floor.R parity. At R's
V(Basque Tier 1), the weights matchSynthto 4.5e-6 (9.6e-5 on main); the test's tolerance goes from1e-3to2e-5. At the defaults (nestedV,seed=0), the pre-RMSPE is 0.0888 againstSynth's 0.0942 and the ATT −0.5842 against −0.5799, with weights 0.8331 (Cataluña) and 0.1669 (Madrid).Methodology references (required if estimator / math changes)
Synth: at an exact-fit tie the reportedWis the active set's vertex whereipoptends to an interior point (REGISTRY Note, above).Validation
test_methodology_synthetic_control.py:_sc_weight_nnlsmeets the simplex QP's KKT conditions (3 factor panels,zeta0 and 0.3) and returns a vertex at an exact-fit tie;_nnls_lawson_hansonequalsscipy.optimize.nnlswhere the optimum is unique, meets KKT on tall, square and wide matrices with a duplicated column, skips a column that rounding lets in, and raises rather than truncate; on factor panels where the budgeted solve stopped, the inner solve is exact,in_space_placebo()keeps every placebo, and the conformal proxy converges.test_prep.py:synthetic_weightis the exact simplex optimum, and stays uniform on a panel with a missing pre-period outcome._truncate_inner_solves(the solve itself cannot fail); the CWZ proxy oracle is held to SLSQP atatol=1e-7with a one-step budget (was1e-2with 200 000 steps); Basque Tier 1 atatol=2e-5(was1e-3); the conformal-interval test on a fine grid (below); Tutorial 25's drift bands; the naming guard'szetaallowlist._sc_weight_nnlsand_nnls_lawson_hansondo not exist on main.test_inner_solve_is_exact_where_a_budgeted_solve_stopped(the budgeted solve does not converge);test_in_space_placebo_keeps_every_placebo_on_a_factor_panel[0]and[1](the treated fit warns "did not converge", and 6 of 19 placebos are dropped);test_conformal_proxy_converges_on_a_factor_panel;test_cwz_proxy_fit_matches_scipy_simplex_ls;test_synthetic_weight_is_the_exact_simplex_optimum(main'ssynthetic_weightobjective 8.9e-11 against the optimum's 3.1e-14); andtest_basque_tier1_custom_v_parityatatol=2e-5.test_top_donor_weights(Tutorial 25's weights) andtest_inner_nonconvergence_warning(the warning no longer names Frank-Wolfe).inner_max_iterno longer acts): no "did not converge" warning, no NaN p-value and no dropped placebo on any of the 120 panels, and the treated pre-RMSPE above the exact optimum by at most 6.7e-16 (relative). The nested count from my comment: no solve excluded fromVranking on any of the 30 fits.DIFF_DIFF_BACKEND=python.pytest -n 4 -m ''(slow tests included) overtest_methodology_synthetic_control.py,test_t25_synthetic_control_policy_drift.py,test_prep.py,test_utils.py,test_methodology_sdid.py,test_diagnostic_report.py,test_business_report.py,test_practitioner.py,test_rust_backend.py,test_aliases.py,test_v4_matrix.py,test_v4_wrapper_shims.py,test_naming_guard.py,test_changelog_fragments.py,test_guides.py,test_doc_snippets.py,test_doc_deps_integrity.py,test_docs_ia.pyandtest_tracking_files.py: 2324 passed, 221 skipped, 2 failed.test_naming_guard.py::test_allowlists_are_reachable: this PR removed the lastzeta=fromsynthetic_control.py, which left itszetaallowlist entry with nothing to cover. The entry is removed; the docs, changelog, prep, guides and naming-guard files then pass (899 passed, 7 skipped).test_rust_backend.py::TestClusterVcovDeterminism::test_raw_kernel_noncontiguous_ids_bit_identicalfails on main too here: that class has noskipif(not HAS_RUST_BACKEND)and importsdiff_diff._rust_backend, which is not built here.ruff check,black --checkandmypyon the changed modules: clean.changelog_compile.py check: OK.synthetic_weighttests (50, selected with-k "t25 or nnls or lawson or exact or synthetic_weight or factor_panel"over the three files) pass under SciPy 1.10.1, 1.11.4, 1.12.0, 1.13.1, 1.14.1, 1.15.3, 1.16.3 (Python 3.11) and 1.17.1 (Python 3.12), NumPy as in the table above.Security / privacy
Changelog
changelog.d/20261006-sc-exact-simplex-weights.md,### Fixedand### Behavioral Changes. It states which results move and by how much (Tutorial 25, Basque, the conformal coverage), the vertex at ties, and thatinner_max_iter/inner_min_decreaseno longer act.