Jefferson Pruett.

Survey methods

Matched Sample Accuracy

On mentorship, how surveys pull the strings behind the world, and taking a closer look at a new sampling method quietly taking over the industry.

A brief note on mentorship

Jon Krosnick was my main undergrad mentor, and this project grew out of his Political Psychology Research Group. In college, between being an instructor at the outdoor center and an ethical deliberation fellow, I interacted a fair amount with underclassmen. Inevitably, they asked a lot of questions to the effect of: how does one actually “make it” at Stanford? I took that to mean: how does one navigate the endless labyrinth of doors that the university can open for you, explore those that interest you, ultimately walk through one, and find personal and intellectual growth by the end of your four years? Reflecting on my own experience, and even more so observing what differentiated my friends from each other, I came to believe that the difference between those who experienced growth, perhaps even transformation, and those who went through the motions and came out the other end wondering what the point was, was that the former found a mentor in the fullest sense of the word. I believe this because of Jon.

With Jon at my Stanford thesis medal ceremony.
With Jon at my Stanford thesis medal ceremony.

Under his often intense tutelage, I learned what it meant to be a legitimate practitioner of open science, and that obsession and relentlessness had a place in academia alongside kindness.

He lent me his “glasses,” so to speak. When you spend a career looking at the world in a particular way, even a niche field begins to open itself up to you, exposing how it pulls the strings on the world around you. For me, that was survey work. A good survey is deceptively difficult to carry out, and while everyone knows what a survey is, few realize how many “facts” are derived from careful statistical models drawing on survey data with known limitations and assumptions. The strings started to pull on why the polls were so bad in 2016, why half of what you read in the Times isn’t quite what it looks like, and the fact that GDP, inflation, and even the number of people incarcerated in the U.S. are all rooted in a survey.

Error in matched samples

My primary project with the lab centered on a careful analysis of one of the largest datasets in political science, and a poster child for a new and rapidly popularizing kind of survey that uses a technique called “sample matching.”

For most of the 20th century, good survey research meant probability sampling: pick respondents at random, with known odds of selection, and the math of inference follows. The internet broke that model economically. Random-digit-dial response rates collapsed, and a wave of online “nonprobability” panels, people who opt in rather than get selected, offered faster, cheaper data instead. The catch, replicated across dozens of studies, is that opt-in samples are usually less accurate.

One firm, YouGov, seems to have found a way around that tradeoff. YouGov’s method, called sample matching, builds a synthetic sample by drawing a random target sample from a real government survey, then finding opt-in panelists who most closely resemble each person in that target sample on a wide set of characteristics. The pitch is that if the matching variables are rich enough, the matched sample behaves statistically like the random one it’s standing in for.

In English, that means a researcher might take a sample of 100 randomly selected people from the Census. Person 89 out of 100 might be 25 years old, white, born in the Northeast, Jewish, and in the 60-70% decile for income. It would be impossible to go find that exact person and ask him what he thought of Trump’s politics, economic policy, or whatever else. So instead, YouGov accrues people from all over the world, keeping tabs on their demographic characteristics. YouGov might go into its database and find someone who, on paper, looks a lot like person #89. Maybe he’s actually 28 and in a lower income bracket, but he’s still white, still Jewish, still from the Northeast. The assumption is that if you ask him what he thinks of Trump’s economic policy, the answer will probably be similar to that of person #89, who you sampled out of the Census.

The Cooperative Election Study, which is administered by YouGov, is one of the most widely used political surveys in academic research and uses this method. A well-known 2016 Pew study found YouGov beating not just other opt-in panels but some probability samples outright. This was kind of like violating entropy. It’s a tenet of survey science that probability samples do better, so this report made waves and raised eyebrows about this method.

What we’re testing, and why it matters

I sought to check whether that reputation holds up at scale, using nine waves of the CES, 2006 through 2022, checked against real external benchmarks like Census demographic data, verified turnout records, and certified election returns. The design choice that matters most is separating variables the CES explicitly builds its statistical weights around from everything else. All surveys use some form of “weight.” This accounts for the fact that, for example, you might have a sample with 60% women and 40% men, but of course we know the split is more like 51/49. So to account for this, a survey might weight the male respondents more, to make them act like 49% of the sample, rather than 40%, if the goal is to say something about how men and women differ in aggregate.

In the above example, we were weighting to a known population proportion. But we could also weight to something like an election outcome. If we know how many people voted for a given president in 2020, we can use that to do our weighting. When analyzing the CES, we had to carefully disentangle which variables were used to construct weights, since error would be mechanistically constrained on these variables, relative to variables not used in weighting. The latter ultimately reflect the performance of sample matching, rather than also mixing in the effects of weighting.

The paper is currently under review, so I’m keeping the specifics light here until that process finishes. What I can say is that the CES isn’t a niche product. It underpins a large share of published political science research on voting behavior, and its accuracy is frequently assumed rather than checked, in part because of results like that 2016 Pew study. Whether sample matching actually closes the gap between opt-in and random samples, or whether that reputation is doing more work than the method itself, has real implications for how much weight researchers should put on findings drawn from it.