Showing posts with label 2020 Elections. Show all posts
Showing posts with label 2020 Elections. Show all posts

Sunday, May 16, 2021

Review: Kalman Filtered Senate Polls

Previously, we discussed Kalman filtering the polls. We will examine how well such a filter performed in the 2020 US Senate races.

Implementation Details

The basic details of the Kalman filter may be found in the previous blog post, here I just review some of the decisions I've made when implementing it in practice:

Polling Date. We treat the date for a poll as the midpoint between its start and end dates. To be clear, we are truncating timestamps to dates.

Polls are 4-vectors. We treat a poll as giving 4 data points: the percentage support for the Democratic candidate, the Republican candidate, all third-party candidates, and the undecided voters.

Third Parties. I treated all third parties as a single candidate. For polls which do not ask about third-party voters, we treat them as the margin of error.

Pooling the Polls. Following Jackman's "Pooling the Polls", we take a precision-weighted average of polls concluding on the same day. This amounts to, for a single poll released on a given end date, renaming the variables and computing the covariance and precision matrices. For multiple polls, this amounts to taking the covariance matrix, inverting it to obtain the precision matrix, multiplying the associated polling result as a 4-vector, weighting the result by the polling size, then multipling the resulting sum by the matrix inverse of the sum of precision matrices. This gives us the "effective responses" for each candidate.

Undecided voters. If a poll ignores undecided voters, we similarly treat them as the margin of error.

Criteria for Assessing Estimates

We will look at the polls at two points of time: at the beginning of October, and a week before the election (i.e., October 27, 2020). The reason for this decision being, we're interested in whether the Democratic candidate could have acted to "course correct" before the election.

The assessment will compare the Democratic candidate's polling average with the votes cast on election day. This is because a lot of Republican supporters are nervous about publicly backing the Republican candidate, and tend to be swept into the "undecided" bucket.

We note that the estimates the Kalman filter produces is a multivariate normal distribution. Since we're interested in the Democratic candidate's performance, we can turn it into a univariate normal distribution. The 95% margin is given in parentheses around the estimates.

Further, we restrict focus to competitive senate seats. Inside Elections considered the following races as "competitive" [i.e., tilt or toss-up]: Arizona, Georgia, Iowa, Kansas, Maine, Montana, North Carolina, South Carolina.

Conclusion: The only state where Kalman filtering the polls produces significantly different results than the observed vote percentage is Maine, where Susan Collins overperformed (or Sarah Gideon drastically underperformed). All other senate race results coincide, within a prescribed margin of error, with the Kalman filtered polling.

Arizona

We plot the polling average as of October 1st:

The Polling averages as of October 27th:

Candidate Sept. 28 Oct. 27 Vote Percent
Mark Kelley (D)46.8% (±4.32%)48.52% (±4.22%)51.16%
Martha McSally (R)38.4%43.37%48.81%
Third Party4.71%3.42%0.03%
Undecided10.0%4.6%--

Georgia Regular

Jon Ossoff faced an uphill battle, but managed to pull off an unexpected victory.

Candidate Oct. 1 Oct. 27 Runoff Vote Percent
Jon Ossoff (D)41.8% (±3.49%)44.87% (±3.40%)47.9%
David Perdue (R)46.8%42.57%49.7%
Third Party2.95%4.08%2.4%
Undecided8.45%8.47%--

Iowa

The polls were fairly accurate for Ms Greenfield, with Kalman filtering at least. We should note the large number of undecideds in the polls reflect the surprisingly large noise.

Candidate Sept. 26 Oct. 25 Vote Percent
Theresa Greenfield (D)47.1% (±4.51%)45.63% (±3.82%)45.15%
Joni Ernst (R)44.0%46.10%51.74%
Third Party1.39%2.89%3.11%
Undecided7.44%5.39%--

Maine

Senator Susan Collins remained stable in the polls around 41% until the very end of October, whereas her Democratic challenger fluctuated between 40% and 50%.

Candidate Sept. 29 Oct. 25 Vote Percent
Sara Gideon (D)45.69% (±6.05%)50.89% (±5.27%)42.39%
Susan Collins (R)42.17%49.0%50.98%
Third Party0.72%0.04989734%6.63%
Undecided11.42%0.06070465%--

Montana

The polls in Montana were far more stable for Steve Bullock.

Candidate Oct. 2 Oct. 26 Vote Percent
Steve Bullock (D)46.01% (±4.3%)46.38% (±4.21%)45.0%
Steve Daines (R)45.40%45.72%55.0%
Third Party3.66%3.31%0%
Undecided4.93%4.60%--

North Carolina

We plot the polling average as of October 1st:

The polling average fluctuates wildly in October, which is hard to see in the plots I could produce, since it's so wild. Remember, Cal Cunningham landed in hot water with an extramarital affair. We see the plot produced October 30, 2020, reflects fluctuation ranges between 45% and 48% for Cunningham:

Candidate Sept. 26 Oct. 27 Vote Percent
Cal Cunningham (D)50.36% (±4%)47.0% (±2.8%)46.9%
Thom Tillis (R)39.30%46.5%48.7%
Third Party4.19%1.42%4.4%
Undecided6.15%5.12%--

South Carolina

It's not clear what caused Sen Graham's support to increase over time, but I suspect it's that undecideds "came back home" to Sen Graham (for whatever reason). Support for Jaime Harrison fluctuated around 44% (with a standard deviation of about ±2%) throughout October, just observing the Kalman filtered results.

Candidate Oct. 2 Oct. 26 Vote Percent
Jaime Harrison (D)44.85% (± 3.9%)41.62% (± 3.6%)44.17%
Lindsey Graham (R)44.82%49.89%54.44%
Third Party3.08%3.38%1.39%
Undecided7.25%5.10%--

Saturday, July 18, 2020

The Economist's 2020 Forecast

Every presidential election, forecasts crop up. Some are qualitative, others are quantitative, some scientific, others humorous (e.g., "tossup bot"). Danger lies when either humorous forecasts are mistaken as serious, or when scientific forecasts err in its approach (by accident or by bad design).

I was struck reading this article on forecasting US presidential races, specifically the passage:

[Modelers] will be rolling out their predictions for the first time this year, and they are intent on avoiding mistakes from past election cycles. Morris, the Economist’s forecaster, is one of those entering the field. He has called previous, error-prone predictions “lying to people” and “editorial malpractice.” “We should learn from that,” he says.
(Emphasis added) I agree. But without model checking, structured code, model checking, some degree of design by contract, model checking, unit testing, model checking, reproducibility, etc., how are we to avoid such "editorial malpractice"?

The Economist has been proudly advertising its forecast for the 2020 presidential election. On July 1st, they predicted Biden would win with 90% certainty. Whenever a model, any model, has at least 90% certainty in an outcome, I have a tendency to become suspicious. Unlike most psephological prognostications, The Economist has put the source code online (relevant commit 980c703). With all this, we can try to dispel any misgivings concerning the possibility of this model giving us error-prone predictions...right?

This post will focus mostly about the mathematics and mechanics underpinning The Economist's model. We will iteratively explain how it works in successive refinement, from very broad strokes to more mundane details.

How their model works

I couldn't easily get The Economist's code working (a glaring red flag), so I studied the code and the papers it was based upon. The intellectual history behind the model is rather convoluted, spread across more than a few papers. Instead of tracing this history, I'll give you an overview of the steps, then discuss each step in detail. The model forecasts how support for a candidate changes over time assuming we know the outcome ahead of time.

The high-level description could be summed up in three steps:

Step 1: Tell me the outcome of the election.

Step 2: Fit the trajectory of candidate support over time to match the polling data such that the fitted trajectory results in this desired outcome from step 1.

Step 3: Call this a forecast.

Yes, I know, the input to the model in step 1 is normally what we would expect the forecast to tell us. But hidden away in the paper trail of citations, this model is based on Linzer's work. Linzer cites a paper from Gelman and King from 1993 arguing it's easy to predict the outcome of an election. We've discussed this before: it is easy to game forecasting elections to appear like a superforcaster.

Some may object this is a caricature of the actual model itself. But as Linzer (2013) himself notes (section 3.1):

As shown by Gelman and King (1993), although historical model-based forecasts can help predict where voters’ preferences end up on Election Day, it is not known in advance what path they will take to get there.
This unexamined critical assumption lacks justification. Unsurprisingly, quite a bit has happened in the past 27 years...there's little reason to believe such appeals to authority. If you too doubt this assumption is sound, then any predictions made from a model thus premised is equally as suspect.

In pictures

Step 1, you give me the outcome. We plot the given outcome ("prior") against time:

We transform the y-axis to use log-odds scale rather than probability, for purely technical reasons. We drop a normal distribution around the auxiliary model's "prior probability" which reflects our confidence in the model's estimates. We then randomly pick an outcome according to this probability distribution (the red dot) which will 68% of the time be within a sigma of the prior (red shaded region):

Then we perform a random walk backwards in time until we get to the most recent poll.1Why backwards? Strauss argued in his paper Florida or Ohio? Forecasting Presidential State Outcomes Using Reverse Random Walks this circumvents projection bias, which tends to underestimate the potential magnitude of future shocks. This technique has been more or less accepted, conveyed only through folklore. Since this is a random walk, we shade the zone where the walk would "likely be":


(The random walk is exaggerated, as most of the figures are, for the sake of conveying the intuition. The walk adjusts the trajectory by less than ~0.003 per day, barely perceptible graphically.)

Once we have polling data, we just take this region and bend it around the polling data transformed to log-odds quantities (the "x" marks):

The fitting is handled through Markov Chain Monte Carlo methods, though I suspect there's probably some maximum likelihood method to get similar results.

Now we need to transform our red region back to probability space:

Note: the forecast does not adjust the final prediction (red endpoint) beyond 2% when polling data is added. There are some technical caveats to this, but this is true for the choice of parameters The Economist employs for 2016 and presumably 2020 (2012 narrowly avoids being overly-restrictive). For this reason I explain step 1 as "tell me the outcome" and step 2 works out how to get there. Adding polling data does not impact the forecast on election day.2I discovered this accidentally by trying to make the 2016 forecast reproducible. The skeptical reader can verify this by changing the RUN_DATE and supplying to rstan a fixed seed (and subsets of a fixed initial values). Fixing the initial values, and the seed supplied to rstan produced identical forecasts on election day. Adding more polling data does not substantively adjust the final predictions in the states, but changing the initial guesses will. I suspected this when I compared the forecasts made at the convention as opposed to on election day: to my surprise, they were shifted by less than 2%. The only substantial source of altering the final prediction stems from precisely what was fixed: (1) randomly selecting a different final prediction (red endpoint), and (2) randomly sampling these initial values (which accounts for a fraction of a difference compared to randomly changing the red endpoint).

And we pretend this is a forecast. We do this many times, then average over the paths taken, which will produce a diagram like this. Of course, if the normal distribution has a small standard deviation (is "tightly peaked" around the prior prediction), then the outcome will hug even tighter to the prior prediction...and polling data loses its pulling power. The Economist's normal distribution is alarmingly narrow for 2016, according to Linzer's work. This means that the red dot forecast will be nearly always around the horizontal line for the auxiliary model's forecast, i.e., amounts to small fluctuations around an auxiliary model. Very likely, with the data supplied, it appears The Economist is using an overly narrow normal distribution for its 2020 forecasts (similar to its 2016 forecasts), which means it's window-dressing for the prior prediction.

If this seems like it's still a caricature, you can see for yourself with how The Economist retrodicted the 2016 election.

Step 1: Tell me the Outcome

How do we tell the model what the outcome is? The chain of papers appeals to "the fundamentals" (from economic indicators and similar "macro quantities", we can forecast the outcome). Let this sink in: we abdicate forecasting to an inferior model, to which we will fit data around. The literature and The Economist use the Abramowitz time-for-change mode, a simple linear regression in R pseudo-code:

Incumbent share of popular vote ~ (june approval rating) + (Q2 gdp)
This is then adjusted by how well the Democratic (or Republican) candidate performed in each state to provide rough estimates for each state's election outcome. The Economist claims to use some "secret recipe" on their How do we do it? page, and it does seem like they are using some slight variant (despite having Abramowitz's model in their codebase).

Linzer notes this is an estimate, and we should accommodate a suitably large deviation around these estimates. We do this by taking normally distributed random variables: \[\beta_{iJ}\sim\mathcal{N}(\operatorname{logit}(h_{i}),\sigma_{i}) \tag{1}\] Where "\(h_{i}\)" is the state-level estimates we just constructed, "\(\sigma_{i}\)" is a suitably large standard deviation (Linzer warns against anything smaller than \(1/\sqrt{20}\approx 0.2236\)), and the "\(\beta_{iJ}\)" are just reparametrized estimates on the log-odds scale intuitively representing the vote-share for, say, the Democratic candidate. (The Economist uses parameters \(\sigma_{i}^{2}\approx 0.03920\) or equivalently \(1/\sigma_{i}^{2}\approx25.51\), which overfits the outcomes given in this step. Linzer introduces parameter "\(\tau_{i}=1/\sigma_{i}^{2}\)" and warns explicitly "Values of \(\tau_{i}\gt20\) [i.e., \(1/\sigma_{i}^{2}\gt 20\) or \(\sigma_{i}^{2}\lt 0.05\)] generate credible intervals that are misleadingly narrow." This is not adequately discussed or analysed by The Economist.)

We should be quite explicit stressing the auxiliary model producing the prior forecast uses "the fundamentals", a family of crude models which work under fairly stable circumstances. The unasked question remains is this valid? Even assessing if the assumptions of the fundamentals holds at present, even that, is not discussed or entertained by The Economist. While experiencing a once-in-a-century pandemic, and the greatest economic catastrophe since the Great Depression, I don't think that's quite "stable".3Given the circumstances, the Abramowitz model underpinning The Economist's forecast fails to produce sensible values. For example, it predicts Trump will win −20% of the vote in DC and merely 34% of the vote in Mississippi (the lowest amount a Republican presidential candidate received since 1968, when Gov Wallace won Mississippi on a third party ticket). This should be a red flag, if only because no candidate could receive a negative share of the votes.

The logic here may seem circular: we forecast the future because we don't know what will happen. If we need to know the outcome to perform the forecast, then we wouldn't need to perform a forecast (we'd know what will happen anyways). On the other hand, if we don't know the outcome, then we can't forecast with this approach.

Bayesian analysis avoids this by trying many different outcomes to see if the forecasting method is "stable": if we change the parameters (or swap out the prior), then the forecast varies accordingly. Or, it should. As it stands now, it's impossible to do this analysis with The Economist's code without rewriting it entirely (owing to the fact that the code is a ball-of-mud). Needless to say, doing any form of model checking is intractable with The Economist's model given its code's current state.

Step 2: "Fit" the Model...to whatever we want

The difference between different members of this family of models lies in how we handle \(\beta_{i,j}\). For example, Kremp expands it out at time \(t_{j}\) and state \(s_{i}\) (poor choice of notation) as \[ \beta[s_{i}, t_{j}] = \mu_{a}[t_{j}] + \mu_{b}[s_{i},t_{j}] + \begin{bmatrix}\mbox{pollster}\\\mbox{effect}\end{bmatrix}(poll_{k}) + \dots\tag{2} \] decomposing the term as a sum of nation-level polling effects \(\mu_{a}\), state-level effects \(\mu_{b}\), pollster effects (Kremp's \(\mu_{c}\)), as well as the "house effect" for pollsters, error terms, and so on. The Economist refines this further, and improves upon how polls impact this forecast. But then every family member uses Markov Chain Monte Carlo to fit polling data to match the results given in step 1. It's all kabuki theater around this outcome. The ingenuity of The Economist's model sadly amounts to an obfuscatory facade.

Here's the problem in graphic detail. With historic data, this is the plot of predictions The Economist would purportedly have produced (the back horizontal lines are the prior predicted, the very light horizontal gray is the 50% mark):

Observe the predictions land within a normal distribution with standard deviation ~2.5% of the prior forecast (the black horizontal line). The entire complicated algorithm adjusts the initial prediction slightly based on polling data. (In fact, none of the forecasts differ in outcome compared to the initial prediction, including Ohio.) Now consider the hypothetical situation where Obama had a net −15% approval rating in June and the US experienced −25% Q2 GDP growth. The prior forecast projects a dim prospect for hypothetical Obama's re-election bid:

This isn't to say that the prediction is immune to polling data, but observe the prediction on election day is drawn from the top 5% of the normal distribution around the initial prediction [solid black line]. From a bad prior, we get a bad prediction (e.g., falsely predicting Obama would lose Wisconsin, Pennsylvania, Ohio, New Hampshire, Iowa, and Florida). This is mitigated if we have a sufficiently \(\sigma^{2}\gt0.05\) as this hypothetical demonstrates.

Also worth reiterating, in 2012, the normal distribution around the prior is wider than in 2016: when the same normal distribution is used, the error increases considerably. We can perform the calculations over again to exaggerate the effect:

Take particular note how closely the prediction hugs the guess. Of course, we don't know ex ante how closely the guessed initial prediction matches the outcome, so it's foolish to make claims like, "This model predicts Biden with 91% probability will win in November, therefore it's nothing like 2016." From a bad crow, a bad egg.

How did this perform in 2016?

Another glaring red flag should be this model's performance in 2016. Pierre Kremp implemented this model for 2016, and forecasted a Clinton victory with 90% probability. The Economist's modifications didn't fare much better, a modest improvement at the cost of drastic overfitting.

Is overfitting really that bad? The germane XKCD illustrates its dangers.

The play-by-play forecasts for multiple snapshots across time. Initially, when there's no data (say, on April 1, 2016), it's very nearly the auxiliary model's forecast:


Note: each state has its own differently sized normal distribution (we can be far more confident about, say, California's results than we could about a perennially close state like Ohio or Florida).

Now, by the time of the first convention July 17, 2016, we have more data. How does the forecast do? Was it correct so far? Well, we plot out the same states' predictions with the same parameters:

Barring random noise from slightly different starting points on election day, there's no change (just a very tiny fluctuating random walk) between the last poll and election day. There's no forecast, just an initial guess around the auxiliary model's forecast.

We can then compare the initial forecast to this intermediate forecast to the final forecast:

Compare Wisconsin in this snapshot to the previous two snapshots, and you realize how badly off the forecast was until the day of the election. Even then, The Economist mispredicts what a simple Kalman filter would recognize even a week before the election: Clinton loses Wisconsin.

The reader can observe, compared to the hypothetical Obama 2012 scenario where Q2 GDP growth was −25% and net approval rating −15%, the Clinton snapshots remain remarkably tight to the prior prediction. Why is this? Because the \(\sigma^{2}\lt0.05\) for the Clinton model, but \(\sigma^{2}\gt0.05\) for the Obama model. This is the effect of overfitting: the forecast sticks too closely to the prior until just before the election.

Although difficult to observe close to election day (given how cramped the plots are), the forecast doesn't change on election day more than a percentage point or two: it changes the day prior, sometimes drastically. We could add more polling data, but the only way for the forecast to substantially change is for the random number generator to change the endpoint, or for the Markov Chain Monte Carlo library to use a different seed parameter (different parameters for the random number generator).

Conclusion

When The Economist makes boastful claims like

Mr Comey himself confessed to being so sure of the outcome of the contest that he took unprecedented steps against one candidate (which may have ended up costing her the election). But the statistical model The Economist built to predict presidential elections would not have been so shocked. Run retroactively on the last cycle, it would have given Mr Trump a 27% chance of winning the contest on election day. In July of 2016 it would have given him a 30% shot.
And no one else was skeptical of a Clinton victory in July 2016? We can recall FiveThirtyEight gave Clinton a more modest 49.9% chance of winning in July compared to Trump's 50.1%; and on election day, comparable odds as The Economist claims. FiveThirtyEight had a "worse" Brier score than The Economist (the only metric they're willing to advertise), but The Economist had the worse forecast in July. Inexcusably worse, The Economist begins by assuming the outcome, then fitting the data around that end. What of that "editorial malpractice" that, as Mr Morris of The Economist urged, "We should learn from"?

We should really treat predictions from The Economist's model with the same gravity as a horoscope. The Economist should take a far more measured and humble perspective with its "forecast". It is more than a little ironic that the magazine which makes such confident claims on the basis of a questionable model once wrote, Humility is the most important virtue among forecasters. Although this election may appear to be an easy forecast — Vice President Biden's lead over President Trump appears large in the polls (routinely double digits) except the undecideds routinely poll in double digits (sound familiar from 2016?)4Remember, when looking at the lead one candidate has over another in a poll ("Biden has a +15% lead"), that poll's margin of error should be doubled, due to the rules of arithmetic of Gaussian distributed random variables. This implies the undecideds plus twice the reported margin of error is routinely equal to or greater than the margin Biden leads Trump. This is an underappreciated point, an eerie parallel to 2016. — we may soon learn The Economist exists to make astrologers look professional.

Sunday, May 10, 2020

States to watch for 2020 Presidential Election

Given there are 50 states (plus DC) consisting of some 3,143 counties, which ones are important and worth keeping an eye on for the 2020 presidential election?

I'm just going to summarize what the big three (Inside Elections, Cook Political Report, and Sabato's Crystal Ball) to determine which states are competitive. Competitive ratings include "toss up", "tilt", and "lean" ratings (in a rather convoluted terminology).

The states considered "in play" are:

  1. Arizona (Cook, IE, Sabato: Toss up)
  2. Florida (Cook, IE: Toss up; Sabato: Leans R)
  3. Georgia (Cook, IE, Sabato: Leans R)
  4. Iowa (IE, Sabato: Leans R)
  5. Maine (Cook: Lean D)
  6. Michigan (Cook: Toss Up; IE: Tilt D; Sabato: Lean D)
  7. Minnesota (Cook, IE, Sabato: Lean D)
  8. New Hampshire (Cook, IE, Sabato: Lean D)
  9. Nevada (Sabato: Lean D)
  10. North Carolina (Cook, IE, Sabato: Toss up)
  11. Ohio (Sabato: Lean R)
  12. Pennsylvania (Cook, Sabato: Toss up; IE: Tilt D)
  13. Texas (Cook: Lean R)
  14. Wisconsin (Cook, IE, Sabato: Toss up)

The eight states in bold are the eight I think are worth studying further. Of course, it's six months out, and a lot can happen in just a couple months (like a virus killing 80,011 people). In a couple months, it may turn out new states come into play, or some of these competitive states cease to be competitive, or both.

(Addendum: there are 176 ways for either candidate to win the presidency.)

Monday, June 24, 2019

How is my candidate doing in the polls?

Given the profusion of polls, it is difficult to accurately gauge how well a given candidate is doing. A simple average of poll numbers won't adequately capture momentum (if such a concept exists), and a few averages (say, one of polls done in the past week, another of polls done in the past month) are difficult to parse. We want one, single, simple number.

We fix the candidate we're interested in, and we have polls \(P_{n+1}\) and \(P_{n}\) released at times \(t_{n}\lt t_{n+1}\). Ideally, we should be able to truncate the N polls to the last k without "much loss".

We could take a moving average, something like \[M_{n+1} = \alpha(t_{n}, t_{n+1}) P_{n+1} + (1 - \alpha(t_{n}, t_{n+1}))M_{n}\tag{1}\] where \(P_{n}\) refers to the nth most recent poll released on the date \(t_{n}\), with the initial condition \(M_{1} = P_{1}\) and the function \[\alpha(t_{n}, t_{n+1}) = 1 - \exp\left(-\frac{|t_{n+1} - t_{n}|}{30 W}\right)\tag{2}\] where W is the average of the intervals between polls, and the difference in dates is measured in days. The 30 in the denominator of the exponent reflects 30 days in a month

Exercise 1. Show (1) α will take values between 0 and 1, (2) the larger the α, the quicker it "forgets" older data, (3) "older data" will be forgotten faster [what happens for regularly released polling data? Say, weekly, a new poll is released, what does α look like?].

Weighing Pollsters

If we knew about poll quality, we could add this in as another factor. Suppose we had a function1The codomain is a little ambiguous, we have it here as \(0\lt Q(P)\lt 1\), but either inequality could be weakened to "less than or equal than" conditions. So it could be extended to include \(0\leq Q(P)\lt 1\) or \(0\lt Q(P)\leq 1\) or even \(0\leq Q(P)\leq 1\). \[Q\colon \mathrm{Polls}\to (0,1)\tag{3}\] which gives each poll its quality (higher quality polls are nearer to 1). Then we could modify our function in Eq (2) to be something like \[ \begin{split} \tilde{\alpha}(t_{n}, t_{n+1}, P_{n+1}) &= Q(P_{n})\cdot\alpha(t_{n},t_{n+1}) \\ &= Q(P_{n})\cdot\left(1 - \exp\left(-\frac{|t_{n+1} - t_{n}|}{30 W}\right)\right)\end{split}\tag{4}\] which penalizes "worse polls" from influencing the moving average. (Since worse polls have smaller Q values, which leads to higher \(1 - Q\) values.)

One lazy way to go about this is to use pollster ratings from FiveThirtyEight, discard "F" rated polls, then take the moving average with \(Q(-)\) the familiar grading scheme used in the United States. (Or, more precisely, the midpoint of the interval for the grade.)

Letter grade Percentage Q-value
A+ 97–100% 0.985
A 93–96% 0.945
A− 90–92% 0.91
B+ 87–89% 0.88
B 83–86% 0.845
B− 80–82% 0.81
C+ 77–79% 0.78
C 73–76% 0.745
C- 70–72% 0.71
D+ 67–69% 0.68
D 63–66% 0.645
D- 60–62% 0.61

The other "natural" choices include (a) equidistant spacing in the interval (0, 1] so D- is given the value \(1/13\) all the way to A+ given \(13/13\), or (b) the roots of an orthogonal family of polynomials defined on the interval [0, 1].

Exercise 2. How do the different possible choices of Q-values affect the running average? [Hint: using the table above, is \(\widetilde{\alpha}\leq 0.61\) is an upper bound? Consider different scenarios, good poll numbers from bad polls, bad numbers from good polls.]

Exercise 3. If we assign \(Q(\mathrm{F}) = 0\) as opposed to discarding F-scored polls, how does that affect the weighted running average?

Some computed examples are available on github, but they're what you'd expect.

Monday, May 20, 2019

Candidate Announcements might be Exponential

Looking back at my prediction, we can see why I was off: the assumption (waiting time between announcements follows an exponential distribution) doesn't even hold. Look at the histogram against the expected distribution:

How can we be sure this is different? ...besides looking...

We can use the Kolmogorov-Smirnov test to see if the histogram differs from the expected exponential distribution. The only problem is there's too little data! There are 20 candidates in my data set, I need it closer to 50 for this test to work. So I just duplicate the data, "smearing" it by adding small amounts like 0.000001 or so, and I do this twice to get a total dataset of 60 "intervals".

The null hypothesis is the data follows the exponential distribution, the alternative hypothesis is the data follows some other distribution. The resulting p-value for the test is 0.03019, so we reject the null hypothesis.

This is based on the huge assumption that we can duplicate the data without any problem, which I have severe doubts about. (Addendum: This reasoning, I realized whilst in Maryland a few days after publication, is invalid, though there exists a technique to fabricate data; since, heuristically, around 30 data points are needed to make an inference, it seems the reasoning below is valid.)

In fact, as a sanity test, lets try simulating candidates announcing they are entering the primary, with λ = 20/136. One quick simulation gives us 16 candidates with intervals between announcements:

This doesn't look too far from the real data. If we had tried the Kolmogorov-Smirnov test without fabricating data, we get a p-value of 0.3387, which tells us we fail to reject the hypothesis this data appears to follow an exponential distribution.

So "what's the right answer"? There's sadly not enough data for us to reject the hypothesis (that the intervals between candidate announcements seem to be exponentially distributed), at least using the frequentist hypothesis testing framework.

Also, I discounted a few candidates which FiveThirtyEight has considered "major" like Andrew Yang or John Delaney (though both candidates announced back in 2017, which make them outliers).

I'll have to look for a statistical test which works on around 20 observations, checking if it fits against an exponential distribution. Maybe there's some Bayesian techniques buried away in Gelman somewhere...

To be clear, however, there is no reason to believe candidate announcements are, a priori, exponentially distributed since timing is contingent on when the (potential) candidate thinks other actors are going o announce. Joe Biden said something to the effect of, he's waiting to announce as late as possible because it's part of a strategy he has. But that depends on his calculations of when "the last possible announcement [relative to his adversaries]" which is explicitly dependent on what other people are doing. An exponentially distributed random variable is memoryless: candidates "wouldn't remember" the last time someone announced.

Ostensibly, a waggish critic might argue, this assumption does hold because it's so damn hard to keep track of the last guy or gal who announced their intention to run for President!

But the only way to know a posteriori whether the actual candidates, by accident or by design, appear to announce with intervals which seem to follow an exponential distribution...is to do some statistical analysis.

Data and scratch work is available on GitHub.

Addendum . One moral to take away from this is to not give a single number as a prediction, but the "HDI" (Highest Density Interval, the region defined as centered on the expected value, the lower bound containing 47.5% of the area, and the upper bound containing 47.5% of the area, so the entire region describes events which are 95% probable). This would've given a large spread, lying between roughly 2 hours and 27.8 days, which is far less interesting as a prediction. For the region containing 85% probability, the interval lies between 29 hours and 14.3 days (14 days, 7 hours).

Whether we use Tukey fences or the HDI, the upper bound for any exponential distribution's prediction would be around \((3\pm\varepsilon)/\lambda\) where \(|\varepsilon|\lesssim 0.03\). Since we're typically waiting for an event to occur, we're only really interested in the upper bound (at least, for candidates declaring their intent to run for President).

And after further consideration, the fabrication of data as I have done it above is incorrect, but it is not an invalid technique provided it is done correctly. Maybe that's a topic for a future post...

Sunday, April 14, 2019

Next Candidate to Announce will be Monday or Tuesday

So, candidates are forming exploratory committees, few are formally announcing. Keeping track of the date an exploratory committee was announced or (if the candidate just filed without forming a committee) the date of FEC filing. The data accumulated so far:

Candidate Announced
Elizabeth Warren December 31, 2018[cnn]
Tulsi Gabbard January 11, 2019[fec]
Julian Castro January 12, 2019[bloomberg.com]
Kirsten Gillibrand January 15, 2019[nytimes]
Kamala Harris January 21, 2019[fec]
Pete Buttigieg January 23, 2019[politico]
Cory Booker February 1, 2019[fec]
Amy Klobuchar February 10, 2019
Bernie Sanders February 19, 2019[fec]
Jay Inslee March 1, 2019[fec]
John Hickenlooper March 4, 2019[fec]
Beto O'Rourke March 14, 2019[fec]
Mike Gravel March 19, 2019[nbc]
Tim Ryan April 4, 2019[nbc]
Eric Swalwell April 8, 2019[fec]

I goofed on thinking Ojeda was the start of the primary process, Senator Warren seems like the opening candidate.

The questions that spring to mind include:

  1. How many people will qualify for the June debates?
  2. Are the Democratic candidates similar in announcement behaviour as past Republican candidates?
  3. When will the next announcement be?

I will dig through the first two questions in a future blog post, but the third question is time sensitive. The short answer is we can expect the next candidate to emerge Monday night or Tuesday morning, the exact probability distribution is plotted below:

This is based on the Jeffreys prior for predicting the exponential distribution using the prior candidates in this cycle as the data points (c.f., [wikipedia]). The expected number of days after the Swalwell's announcement is 98/13 ≈ 7.53846.

I'll "show my work" in a future blog post, and include a "DIY" equation to predict the next announcement based on N candidates already announced and d (the number of days between the latest candidate and the first, Elizabeth Warren's announcement date).

Addendum . Representative Seth Moulton (D-MA, 6) was the next candidate to announce he was running for president, throwing his hat in the ring on April 22, 2019: a full week after expected. The probability of this happening, with the Bayesian posterior used here, is approximately 1.927626% whereas the maximum likelihood estimate would give 1.93336%; however we cut it, it's around 1.93% probability. (I forgot to publish this addendum on the date I wrote it, it was saved as a draft for about a month.)

(After further thought, Moulton declared \(2/\lambda\) days after the previous candidate, with \(\lambda=13/98\approx 1/7.5\). The probability of a candidate declaring after twice the expected value in an exponentially distributed model is \(\Pr(x\geq 2/\lambda)\approx 0.13533\) which isn't unreasonable. An event occurring with 13.5% probability is roughly the same odds as getting 3 heads in a row with a fair coin.)

Tuesday, November 6, 2018

When will the 2020 Election "Start"?

Pundits are already speculating about who will run for president on the Democratic ticket for 2020, which begs the question: what does the data suggest when the first person will declare? How many candidates can we expect? And what is the rate at which candidates will enter the race?

Shut up and tell me the answers

Assuming the Democratic race will be like the Republican 2012 and 2016 race, apparently candidates declaring their entrance follows a Poisson Distribution, as the interval between announcements follows a Exponential Distribution; with 95% probability, candidates announce their intent to run anywhere between 7.62 days on the low end, and 17.54 on the high end, with an expected rate of about 11 days between each announcement.

Consequently, since the first serious contender announced Sunday 11 November 2018, we can expect, with 95% probability, anywhere between 21 to 48.55 candidates to make the announcement, with likelihood maximized at 33.6 candidates to be the Democratic nominee in the 2020 cycle.

Rate at which Candidates Declare

According to Wikipedia, the 2012 Republican candidates declared in the following order:

CandidateDate Declared
Newt Gingrich
Ron Paul
Herman Cain
Mitt Romney
Rick Santorum
Jon Huntsman, Jr.
Michele Bachmann
Rick Perry

Similarly for the 2016 Republican primary race, again from wikipedia, we have the following:

CandidateDate Declared
Ted Cruz
Rand Paul
Marco Rubio
Ben Carson
Carly Fiorina
Mike Huckabee
Rick Santorum
George Pataki
Lindsey Graham
Rick Perry
Jeb Bush
Bobby Jindal
Chris Christie

The interval between each candidate declaring his or her candidacy for Republicans in 2012 is (in days): 8, 12, 4, 15, 6, 47. This has a mean of 15.33333 days, and a variance of 256.66666 days. (Observe that the variance is approximately the square of the mean.)

The intervals between each candidate for Republicans in 2016 is (in days): 6, 21, 1, 22, 1, 4, 3, 11, 9, 6. This has a mean of 8.4 days, and a variance of 57.822222 days (the fractional part is 37/45). (Again, observe how the mean squared is approximately the variance.)

When we combine this data together, the concatenated dataset has a mean of 11 days, and a variance of 132.26666 days (observe the square of the mean is approximately the variance).

This property (the variance is approximately the square of the mean) supports the hypothesis that this dataset is described by a Random Variable following an Exponential Distribution with its rate parameter approximately 1/λ ≈ 11 days.

Well, really, 95% of the posterior distribution lies between 0.05702248 < λ < 0.1312337 and peaks at 1/λ ≈ 11 days; the high density interval thus described is shaded in blue in the following plot:

When first candidate will declare

We should probably ask the question How many days before election day November 3, 2020 will the first candidate announce his or her bid for presidency?

The data for the Republican candidates in 2016, dataset (of days before the election the candidate announced) is: 497, 503, 512, 523, 526, 530, 531, 553, 554, 554, 575, 581, 596. Observe this occurred within a span of 199 days.

Similarly, the data for the 2012 election: 451, 498, 504, 519, 523, 535, 543, 545. Also observe this occurred within the span of 94 days (roughly half the length of the 2016 election).

Observe that the 2016 election had 5 candidates declare earlier than the earliest nominee in 2012, so it stands to reason to suppose that the greater the number of candidates, the earlier the first bid, and vice-versa:

Herd-Size Conjecture: the number of candidates is directly proportional to the earliest candidate's declaration date (as measured by days before the election).

(This is not an unreasonable conjecture, since candidates declare at intervals which are described by an exponential distribution, and there is a hard deadline to declare your candidacy.)

Assuming every candidate wants to run in every primary, South Carolina requires filing for primary candidates by September 30, 2019 (assuming it is like the 2016 primary — the deadlines for the 2016 primaries may be found here). This is 400 days before the election, everyone must file before then.

There is actually a fairly decent correlation between the total funds raised, total spent, and total left on-hand (just add up the quarterly books) and the length of a primary campaign. For the nominee, I consider the "end date" to be election day.

I had to work with 2012 data because it's far neater than the 2016 data. The R2 = 0.8546118, and R2adj = 0.6607609; the model is:

(number of days) 
= 115.53788314621659
  - 11.278698441397473×(total raised in $Mn)
  + 22.78101502109625×(total spent in $Mn)
  - 10.49938850575515×(total left on-hand in $Mn)

Supposing this model also holds for the Democrats, all we have to do is estimate how much money is "out there" for candidates. For the Republicans in 2012, all candidates raised a total sum of $337,615,860

Corollary: If the Herd-Size Conjecture holds, then number of candidates is directly proportional to total funds raised by the candidates.

Proof sketch: There are several steps in the proof.

  1. Since the length of a campaign is directly proportional to the funds raised, and the Herd-Size Conjecture says the earliest candidate's declaration (i.e., the start of the campaign) is directly proportional to the number of candidates, it follows the candidate's declaration date is proportional to the funds raised.
  2. Since the candidates declare at a fairly steady rate following an exponential distribution, it follows the earlier the first candidate declares, the greater the number of candidates will declare.
  3. The greater the funds raised, the earlier the candidate's start date tends to be (by step 1), and hence the greater the number of candidates (by step 2). ■

Of course, this just punts the problem (of determining who will declare first) to the much more difficult problem of how much money will be in the Democratic 2020 primaries, and how will it be partitioned.

Something to consider is that the amount may be approximately the total amount of money raised by the party in the House from the previous midterm election (2010 Republicans raised a total of $353Mn in the House, according to OpenSecrets). If this is a good approximation, then there will be roughly $649Mn raised by the Democratic candidates in the primary, and the total amount spent will vary between $421Mn to $649Mn (topic for future post!). But if $443Mn is spent, we could expect to see primary announcements as early as December 4th, 2018.