← Back to the post
Interactive data dictionary

Proposition 99 and cigarette sales in 39 US states (synthetic control tutorial)

This panel of Abadie, Diamond, and Hainmueller (2010) records cigarette sales and four covariates for 1970–2000. It supplies the data behind every estimate in the Python synthetic control tutorial.

39
states (1 treated, 38 donors)
31
years, 1970–2000
1,209
state-year rows
−18.98
ATT in packs per capita per year

Downloads

Each dataset is available as a labeled Stata .dta and its source file.

⇩ Download all data (ZIP)stata_codebook.do

DatasetGrainRowsStataSource
smoking_scstate × year (balanced panel)1,209 × 7smoking_sc.dtasmoking_sc.csv

Run stata_codebook.do in Stata once to attach long-form per-variable notes to the .dta files.

Load directly in code

Every file loads straight from GitHub (raw URLs). Swap the file name to load any dataset.

Stata

* Stata 14+ : `use` reads an https URL directly
global BASE "https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_sc101/data/"
use "${BASE}smoking_sc.dta", clear
describe
notes

Python

!pip install -q pyreadstat
import pandas as pd
BASE = "https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_sc101/data/"
df = pd.read_stata(BASE + "smoking_sc.dta")

# load every dataset at once
files = ["smoking_sc"]
data = {f: pd.read_stata(BASE + f + ".dta") for f in files}

# pyreadstat (richest metadata) reads LOCAL files -> download first
import pyreadstat, urllib.request
urllib.request.urlretrieve(BASE + "smoking_sc.dta", "smoking_sc.dta")
df, meta = pyreadstat.read_dta("smoking_sc.dta")

Copy and paste this snippet into an empty Google Colab notebook: https://colab.research.google.com/notebooks/empty.ipynb

R

# R : haven::read_dta auto-downloads an https URL
library(haven)
BASE <- "https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_sc101/data/"
df <- read_dta(paste0(BASE, "smoking_sc.dta"))

Overview & sources

Proposition 99 raised the cigarette tax in California by 25 cents per pack from January 1989 and funded anti-smoking education. Abadie, Diamond, and Hainmueller (2010), hereafter ADH (2010), evaluate it with a panel of 39 US states from 1970 to 2000. Each row is one state in one year, so the panel has 1,209 rows. California is the treated state, and the other 38 states form the donor pool. The outcome cigsale measures cigarette sales per capita, in packs. Four covariates describe each state: log GDP per capita, the share of the population aged 15–24, the retail price of cigarettes, and beer consumption per capita. The Python tutorial builds a synthetic California from these variables. It estimates that the program reduced sales by 18.98 packs per capita per year over 1989–2000.

One balanced panel. The file smoking_sc has 39 states × 31 years = 1,209 rows, keyed by state × year and sorted by state and year. The sample omits four states that began large tobacco control programs in 1989–2000 (Arizona, Florida, Massachusetts, and Oregon). It also omits seven states that raised cigarette taxes by 50 cents or more over the same period, as well as the District of Columbia (Section 3.2 of ADH 2010). The original Stata file stores state as a numeric code with value labels in alphabetical order, so California has code 3. The CSV copy stores the state name as text instead, and so does the .dta file generated from it in this folder.

Data sources

SourceProvidesReference / URL
Abadie, Diamond, and Hainmueller (2010)The case study and the state panel. Appendix A names the original sources: Orzechowski and Walker (2005) for sales and prices, the Bureau of the Census for income and the age share, and the Beer Institute for beer consumption.Journal of the American Statistical Association, 105(490), 493–505. https://doi.org/10.1198/jasa.2009.ap08746
QuaRCS-lab data-openThe original Stata file smoking_sc.dta, labeled Tobacco Sales in 39 US States, which the post and its Stata edition loadhttps://github.com/quarcs-lab/data-open/raw/master/isds/smoking_sc.dta
This post (CSV copy)smoking_sc.csv, with the state name as text and the values of the original .dta file at full double precision (89.8 appears as 89.80000305175781). The post loads this copy first and falls back to the original file online.Mendez, C. (2026). https://carlos-mendez.org/post/python_sc101/

Cite this data

Please cite this dataset as follows.

APA

Mendez, C. (2026). Proposition 99 and cigarette sales in 39 US states (synthetic control tutorial) [Data set]. https://carlos-mendez.org/post/python_sc101/

Abadie, A., Diamond, A., & Hainmueller, J. (2010). Synthetic control methods for comparative case studies: Estimating the effect of California's tobacco control program. Journal of the American Statistical Association, 105(490), 493–505. https://doi.org/10.1198/jasa.2009.ap08746

BibTeX

@misc{mendez2026pythonsc101,
  author       = {Mendez, Carlos},
  title        = {Proposition 99 and cigarette sales in 39 US states (synthetic control tutorial)},
  year         = {2026},
  howpublished = {\url{https://carlos-mendez.org/post/python_sc101/}},
  note         = {Data set}
}

@article{abadie2010synthetic,
  author  = {Abadie, Alberto and Diamond, Alexis and Hainmueller, Jens},
  title   = {Synthetic Control Methods for Comparative Case Studies: Estimating the Effect of {California's} Tobacco Control Program},
  journal = {Journal of the American Statistical Association},
  volume  = {105}, number = {490}, pages = {493--505}, year = {2010},
  doi     = {10.1198/jasa.2009.ap08746}
}

Variable explorer search & filter all 7 variables

Type to filter by name or label, or use the chips to filter by type. Each row shows a mini distribution. Click a header to sort.

VariableTypeDistributionLabelDefinitionUnitsIn filesSource
age15to24#continuousmin 0.129 | median 0.178 | max 0.204Share of the population aged 15–24, a fraction (covariate)Share of the state population aged 15 to 24, stored as a fraction between 0.129 and 0.204 (mean 0.175). The original label calls it a percent, and Table 1 of ADH (2010) reports it in percent, so multiply it by 100 to compare.fraction (0–1)smoking_scUS Census Bureau, via ADH (2010) and quarcs-lab/data-open
beer#continuousmin 2.5 | median 23.3 | max 40.4Beer consumption per capita, in gallons (covariate)Per capita consumption of malt beverages, in gallons. The data start in 1984, so the 1980–1988 average in the post rests on 1984–1988 alone.gallons per capitasmoking_scBeer Institute, via ADH (2010) and quarcs-lab/data-open
cigsale#continuousmin 40.7 | median 116 | max 296Cigarette sales per capita, in packs (outcome)Annual cigarette sales per capita, in packs, and the outcome of the analysis. The series divides the tax-paid sales of cigarette packs in a state by its population (Appendix A of ADH 2010).packs per capita (annual)smoking_scOrzechowski and Walker (2005), via ADH (2010) and quarcs-lab/data-open
lnincome#continuousmin 9.4 | median 9.86 | max 10.5Log of state GDP per capita (covariate)Natural logarithm of GDP per capita in the state, as the original .dta label and Table 1 of ADH (2010) describe it. The text and Appendix A of ADH (2010) call it per capita state personal income (logged), converted to 1997 dollars with the Consumer Price Index.log of 1997 dollars per personsmoking_scBureau of the Census, United States Statistical Abstract, via ADH (2010) and quarcs-lab/data-open
retprice#continuousmin 27.3 | median 95.5 | max 351Retail price of cigarettes, in cents per pack (covariate)Average retail price of a pack of cigarettes, in cents, including state sales taxes where applicable. Appendix A of ADH (2010) converts income to 1997 dollars but mentions no such conversion for this price.cents per packsmoking_scOrzechowski and Walker (2005), via ADH (2010) and quarcs-lab/data-open
state#identifiern/aState name (39 US states)Name of the US state. California is the treated state, and the other 38 states form the donor pool.smoking_scADH (2010), via quarcs-lab/data-open
year#yearn/aYear (1970–2000)Calendar year of the observation. Proposition 99 takes effect in 1989, so 1970–1988 is the pre-treatment period and 1989–2000 the post-treatment period.calendar yearsmoking_scADH (2010), via quarcs-lab/data-open

Cross-file variable index

Which file each variable appears in (● = present).

Variablesmoking_sc
age15to24●
beer●
cigsale●
lnincome●
retprice●
state●
year●

Construction & formulas

How the post builds the predictors

What the post estimates from this file

The datasets

Switch datasets with the tabs. Each shows the full variable dictionary plus a sortable statistics table with mini distributions and data coverage.

expand to search (Ctrl/⌘+F) or print across all datasets

state × year (balanced panel)  1,209 × 7 · 1970–2000 (31 years) · 39 US states (California and 38 donor states)

Panel key: state × year · The post uses this file for every estimate: the baseline fit, the placebo tests, the leave-one-out refits, and the estimator tour.

Variable dictionary

VariableLabelDefinitionConstructionUnitsSourceCoverage
state identifierState name (39 US states)Name of the US state. California is the treated state, and the other 38 states form the donor pool.The original .dta file stores a numeric code with value labels in alphabetical order (1 = Alabama, 3 = California, 39 = Wyoming). The CSV keeps the label text.ADH (2010), via quarcs-lab/data-open39 states in every year
year yearYear (1970–2000)Calendar year of the observation. Proposition 99 takes effect in 1989, so 1970–1988 is the pre-treatment period and 1989–2000 the post-treatment period.Stored as a float in the original .dta file and as an integer in the CSV.calendar yearADH (2010), via quarcs-lab/data-open31 years for every state
cigsale continuousCigarette sales per capita, in packs (outcome)Annual cigarette sales per capita, in packs, and the outcome of the analysis. The series divides the tax-paid sales of cigarette packs in a state by its population (Appendix A of ADH 2010).Original .dta label: cigarette sale per capita (in packs). The post compares California with its synthetic control on this variable in every year.packs per capita (annual)Orzechowski and Walker (2005), via ADH (2010) and quarcs-lab/data-open1970–2000, all 1,209 rows
lnincome continuousLog of state GDP per capita (covariate)Natural logarithm of GDP per capita in the state, as the original .dta label and Table 1 of ADH (2010) describe it. The text and Appendix A of ADH (2010) call it per capita state personal income (logged), converted to 1997 dollars with the Consumer Price Index.Original .dta label: log state per capita gdp. The post averages it over 1980–1988 as a predictor.log of 1997 dollars per personBureau of the Census, United States Statistical Abstract, via ADH (2010) and quarcs-lab/data-open1972–1997 (1,014 of 1,209 rows)
beer continuousBeer consumption per capita, in gallons (covariate)Per capita consumption of malt beverages, in gallons. The data start in 1984, so the 1980–1988 average in the post rests on 1984–1988 alone.Original .dta label: beer consumption per capita. ADH (2010) also average it over 1984–1988 (notes to Table 1).gallons per capitaBeer Institute, via ADH (2010) and quarcs-lab/data-open1984–1997 (546 of 1,209 rows)
age15to24 continuousShare of the population aged 15–24, a fraction (covariate)Share of the state population aged 15 to 24, stored as a fraction between 0.129 and 0.204 (mean 0.175). The original label calls it a percent, and Table 1 of ADH (2010) reports it in percent, so multiply it by 100 to compare.Original .dta label: percent of state population aged 15-24 years. The post averages it over 1980–1988 as a predictor.fraction (0–1)US Census Bureau, via ADH (2010) and quarcs-lab/data-open1970–1990 (819 of 1,209 rows)
retprice continuousRetail price of cigarettes, in cents per pack (covariate)Average retail price of a pack of cigarettes, in cents, including state sales taxes where applicable. Appendix A of ADH (2010) converts income to 1997 dollars but mentions no such conversion for this price.Original .dta label: retail price of cigarettes. The post averages it over 1980–1988 as a predictor.cents per packOrzechowski and Walker (2005), via ADH (2010) and quarcs-lab/data-open1970–2000, all 1,209 rows

Distribution & statistics (click a header to sort)

VariableDistributionCoverageNDistinctMinMeanMedianMaxSD
staten/a100%1,20939n/an/an/an/an/a
yearn/a100%1,2093119701985.0198520008.95
cigsalemin 40.7 | median 116 | max 296100%1,20970340.70118.9116.3296.232.77
lnincomemin 9.4 | median 9.86 | max 10.584%1,0141,0149.409.869.8610.490.171
beermin 2.5 | median 23.3 | max 40.445%5461452.5023.4323.3040.404.22
age15to24min 0.129 | median 0.178 | max 0.20468%8198190.1290.1750.1780.2040.015
retpricemin 27.3 | median 95.5 | max 351100%1,20984927.30108.395.50351.264.38

Known limitations & caveats