Did Proposition 99 reduce cigarette sales in California?
In January 1989, Proposition 99 raised the cigarette tax in California and funded anti-smoking education. Cigarette sales, however, were already falling across the United States. The decline after 1989 therefore needs a comparison before anyone can credit it to the program. The synthetic control method builds that comparison from a weighted average of other states, chosen to track California before 1989.
This app retells the post in four tabs, and it reads every number from the results file of the post. The tiles below give the answer in brief, and the chart under them shows where it comes from. The glossary at the end of this tab defines each key term in a few sentences.
Observed and synthetic California, 1970–2000
What to look for
- Before 1989, the two paths nearly overlap. After a miss of … packs in 1970, the largest yearly miss is … packs, in ….
- From 1989, the paths separate. The gap is … packs in 1989 and … packs in 2000, when sales sit … percent below the synthetic path.
- A few states build the counterfactual. Only … of the … donor states receive weight, and the next tab shows the recipe.
Donor recipe and gap
Placebo and robustness checks
mlsynth estimator tour
Glossary: open a card when a term is unfamiliar
Synthetic control method (SCM)
The synthetic control method builds a comparison unit from a weighted average of untreated units. The weights make this synthetic unit reproduce the treated unit before the policy. After the policy, the synthetic unit estimates the outcome that the treated unit would have recorded without it. Abadie, Diamond, and Hainmueller (2010) developed the method with the case of Proposition 99.
Donor pool
The donor pool is the set of untreated units that may enter the synthetic control. Here it contains the … states other than California in the data. Abadie, Diamond, and Hainmueller (2010) had already removed states with large tobacco programs or large tax increases. Each donor must stay free of the treatment and of similar policies, or the counterfactual becomes misleading.
Donor weights (the vector W)
Donor weights state how much each donor contributes to synthetic California. They cannot be negative, and they must sum to one, so synthetic California stays inside the range of the donors. The fit of mlsynth gives positive weight to … states, led by … with …. The other … donors receive a weight of zero.
Predictors and predictor weights (the matrix V)
Predictors are the characteristics on which synthetic California must resemble California. The post uses four covariates averaged over 1980–1988 and cigarette sales in 1975, 1980, and 1988. The predictor weights on the diagonal of V set how much a mismatch on each predictor counts. Very different V can select almost the same donor weights, so V is not identified and does not rank the predictors by importance.
Pre-treatment fit (RMSE)
Pre-treatment fit measures how closely synthetic California tracks California before the program. Its usual summary is the root mean squared error (RMSE) of the yearly gaps over 1970–1988. Here the RMSE is … packs per capita, about … percent of mean sales in those years. A close fit is necessary for a credible counterfactual, but it is not sufficient.
Gap and ATT
The gap is the difference between actual and synthetic sales in a given year. The average treatment effect on the treated (ATT) is the mean gap over 1989–2000, the years after the program began. The ATT is … packs per capita per year, a reduction of … percent relative to synthetic sales. It measures the effect on California only, not the effect that the program would have elsewhere.
In-space placebo and MSPE ratio
An in-space placebo test applies the method to every donor as if it had been treated. The mean squared prediction error (MSPE) averages the squared gaps, and the MSPE ratio divides its value after 1989 by its value before 1989. California has the largest ratio of all … states, …. Its permutation p-value is therefore …, the smallest value that … states allow.
In-time placebo and leave-one-out
An in-time placebo moves the start of the program to a year when nothing happened. With a fake start in …, the fake gaps before 1989 average … packs, about one third of the mean gap of … over 1989–2000. A leave-one-out check refits the model once without each donor that receives weight. Across the … refits, the ATT stays between … and … packs.
The donor recipe: which states build synthetic California?
The nested search of mlsynth chooses donor weights that are nonnegative and sum to one. The solution is sparse, because only … of the … donor states receive positive weight. Gold ticks mark the weights that synth2 reports in the Stata log, so the chart doubles as a replication check. Below the weights, a table compares the predictors, and a final chart traces the gap that the recipe implies.
Donor weights: mlsynth bars and Stata ticks
Predictor balance
Predictor weights V: mlsynth and Stata
The two programs put very different weights on the same seven predictors. The fit of mlsynth spreads V almost equally over the age share (…), the retail price (…), and sales in 1975 (…). Stata instead puts … on the age share and … on sales in 1975. Both vectors select almost the same donor weights, so V is not identified, and its entries are not measures of importance.
Actual and synthetic California
What to look for
- Most donors receive zero weight. In total, … of the … donor states receive a weight of zero in both programs.
- The weights repair the balance. Across the seven predictors, the mean absolute percent gap falls from … percent for the donor average to … percent for synthetic California.
- Rounding barely matters. The rounded Stata weights move the ATT from … to … packs, and the largest difference in a donor weight is …, for ….
- V differs, while W agrees. The two programs report very different predictor weights, yet they select almost the same recipe.
Can we break the result? Placebo tests and robustness checks
Every synthetic control misses by some amount, so a large gap alone does not prove an effect. The in-space placebo test refits the model with each state as if it had been treated and asks how unusual California is. Two further checks move the start date and remove one donor at a time. Each section below has its own controls, and every number updates with them.
In-space placebo test
The test refits the model once for every state, each time as if that state had adopted the program. California should stand out if the program mattered. The cutoff of synth2 can drop placebo states whose fit before 1989 is poor, which changes the comparison set.
Gaps of California and the retained placebo states
Pointwise p-values after the program
MSPE ratios of all states
In-time placebo test
The in-time placebo pretends that the program started before 1989. Every predictor is then measured before the fake start, and the gaps between the fake start and 1989 are fake effects. A sound design shows fake gaps that are small relative to the real effect.
Leave-one-out refits
Synthetic California rests on a few donors, so a single donor could drive the result. The leave-one-out check refits the model without each donor that receives weight, one at a time. A robust effect keeps its sign and its rough size in every refit.
What to look for
- California ranks first under every filter. Its MSPE ratio of … stays at rank … with all states, cut(5), cut(2), and cut(1).
- The p-value equals one over the number of retained states. It rises from … with all … states to … with cut(1), because a smaller comparison set has a higher floor.
- The cut removes badly fitted placebos. A large pre-treatment MSPE sits in the denominator of the ratio, so these states usually have small ratios, and removing them does not move California.
- The fake gaps are about one third of the real ones. With a fake start in …, the mean fake gap of … packs is … times the baseline ATT.
- The leave-one-out band stays negative. After 1988, every refit gap lies below … packs, and the gap in 2000 stays between … and ….
One case, four mlsynth estimators
The mlsynth library runs many estimators behind the same configuration dictionary. This tab compares the baseline with three of them, each with a different idea of what a counterfactual should match. All four target the same estimand, the ATT for California over 1989–2000. Tick the boxes to compare their counterfactual paths, and then read the dot plot and the table.
Counterfactual paths
ATT of each estimator
Comparison table
The configuration dictionaries
Every estimator in mlsynth reads one Python dictionary. The dictionary
base holds the data keys, and each estimator adds its own
settings to a copy of it. The baseline adds the predictors, their
windows, and the settings of the nested search.
What to look for
- All four estimates are negative. They range from … for … to … for ….
- SDID shifts the level. It matches the pre-treatment outcomes only up to a constant and adds time weights, and its ATT of … is the least negative of the four.
- PCR uses signed weights. CLUSTERSC spreads its weight over … donors and gives … of them negative weights, so its counterfactual can extrapolate beyond the donors.
- A lower RMSE is not more credible. CLUSTERSC has the lowest pre-treatment RMSE, …. The outcome-only fit also beats the baseline (… against …), yet its placebo test places California at rank … (p = …).
- TWFE lies far from all four. The TWFE reference weights all … donors equally and gives ….