Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Sunday, November 18, 2012

A Measure of Success

538's estimate of state probabilities.
Came up with a way to measure 538's success in predicting the election other than a simple boolean for each state. All predictions are weighted by their certainty. A 50/50 estimate would thus count for nothing (it's worth pointing out that those calling the election a tossup were doing this...), whereas a prediction of 100% would yield a weight of 1.0. Correct predictions are positive, incorrect negative (538 got all 50 states right, plus DC). This weight is then multiplied by a state's electoral votes, summed, and then normalized on the range of the worst score (100% certainty on all, but getting them all wrong: -538) to the best (100% certainty, all correct: 538). Doing this gives 538 a score of .9485. Note that I treated states that award their electoral votes in a split manner as all or nothing for simplicity.

Baseline estimate of state probabilities
To provide more insight, a naive model was used to provide a baseline: all swing states estimated at 50/50 and all non swing states at 100%. This gives a score of 0.8977. This means that 538's estimate gets us 49.66% of the way from the baseline to a perfect forecast. Unfortunately, the data necessary to apply this measure to other forecasts wasn't easily available (and by "easily", I mean "with the amount of effort I was willing to put into this blog post"), but I grabbed some from a couple other sites.

PEC's last estimate map before the election.
Princeton Election Consortium's last posted data was in the form of a color map, so I extracted their estimates by examining the colors and they were clearly not that precise. Applying my measure, they got a score of 0.9429, and were 44.19% of the way from the baseline to a perfect forecast.

Simon Jackman's swing state estimate data.
Simon Jackman over at The Huffington Post had this estimate of swing states probabilities the day before the election. I couldn't find the non swing state opens, so filled in 100% for them, which almost certainly inflated his overall score. It worked out to 0.9544, and was 55.38% of the way from the baseline to perfect.

If anyone has other data (or better versions of the data I did use) I'd be happy to include it in this post. Here is the spreadsheet I used to compute the scores.

Bonus: I edited these maps that have been making the rounds since the election that scale state size according to electoral votes and population respectively to have purple shading that reflects the percentage of Obama and Romney votes. They don't really warrant a post of their own, especially since I didn't create the original maps they're based on.

Size proportional to electoral votes (1 square = 1 vote)
Size proportional to population (original image didn't include HI or AK)



Monday, November 1, 2010

Life Expectancy with Age


As we age our life expectancy increases, as the chance that we might die before that point has been ruled out. By examining 2004 insurance company life tables, I was able to create this rather interesting graph of the effect. You can easily see that the influence of improved health choices and superior genetics begins to give a significant advantage during the early fifties. Of course, except right after birth, the total years remaining to live steadily declines.


Wednesday, October 6, 2010

Binomial Probability Distribution Tree

Math is my best friend.

I frequently work with binomial distributions, and as a visualization aid I created this tree of probability distributions for each of its states up to fourteen trials. Each node represents the beta distribution formed for a given number of successes and failures. At the top is the case with zero data and the probability is spread evenly, as expected. The distributions are shown in white, and the green cup is a positional reference. The red lines lead to the node that adds one additional success to the number of trials and the blue lines similarly lead to the node that adds one additional failure. As you'd expect, the more failures there are, the more the probability distribution crowds to the left and vice versa with increasing successes (notice they have mirror symmetry left to right). You can see the distribution becomes more concentrated as the number of data points increase (the effect is most easily observed straight down the center, where the number of successes equals the number of failures, so the mean stays constant).

Here are a couple more in a different style and varying scales. The last one has the distribution means shown in green.


Thursday, November 5, 2009

The Real World Birthday Problem

Originally posted on my private blog September 19th, 2008.

The birthday problem is a well known mathematical illustration. Basically it asks how many people are needed in one room before two of them have a 50% chance of sharing a birthday. The answer is 23, which apparently is lower than what most people's intuition suggests. That result, however, is based on math that assumes that birthdays occur uniformly throughout the year and that also neglects leap days. I was curious what effect this might have, so I went looking for some birth statistics. The closest I found was a list of U.S. births by day for 1978. Series 1 (dark blue) is the raw information, which as you can see has a weekly cycle with much lower birth rates on weekends (especially Sunday). Unfortunately I needed to remove this effect as every date occurs on every day of the week in the long run. Not having other years to work with, I filtered the data, averaging every day with the births for three days before and after so that every weekday would be included. This is shown as series 2 (pink). Series 3 (yellow) is what the assumption of completely uniform birth rates would look like and is included for comparison. You can see that birth rates are significantly higher in the second half of the year. Unfortunately, 1978 wasn't a leap year, so I averaged the values of February 28th and March 1st and then weighted the leap day to be one quarter as likely (not shown).
Next I had to use this data to compute the probability distribution. I went with a simulation approach, which means I just performed the experiment a lot of times and kept track of how many people it took to get a matching pair of birthdays. Not sure of how many trials were needed to be accurate, I did successively larger runs, starting with ten thousand and ending with a billion, each run being an order of magnitude larger. You can see that the results converge, with the last two curves being right on top of each other.
Finally, here are the cumulative probability distributions of both the standard mathematical model (series 1) and my simulated one using real data (series 2). You can see that the two are basically identical and in fact the maximum error is less than two tenths of one percent. This means that all the effects of non-uniform birthday occurrence average out in practice. If the data had been more dramatically clustered this might not have been the case, but now I know for sure. In case you're wondering, a couple professors have written papers on this very topic, but I wanted to do it for myself as a mental exercise.

Friday, September 11, 2009

News and Opinion Separation & News Accreditation


Like you didn't already know who I had in mind?

It is now common practice for opinion pieces and news stories to be closely juxtaposed and this has the effect of confusing less sophisticated viewers about what is actually fact. There should be strict limits on the placement of the two types of content (disclaimers wouldn't be adequate), ideally limiting them to entirely different programs. Additionally, outlets claiming to present news should have to be certified. On a regular basis polls should be taken about their viewership's knowledge of current events, much as some private studies do now. They'd have to be properly sized and controlled for statistical significance, of course. Any program failing to meet a minimum threshold could not call itself a news program. They could still present any content they like, so it's not really censorship, but rather much more like truth in advertising.

As a further extension, news programs shouldn't be allowed to generate revenue in the usual manner. Instead, any compensation received should be a function of not only audience size, but also accuracy measures. Any program labelled as news should be required to use this system. The gating accuracy measures mentioned previously would still be needed, as there can be motives other than profit for disseminating false information.

Saturday, July 11, 2009

Weighted Song Playing Script

Note that this post has superseded by a new, smart playlist based approach. You can check it out here, where it's available for download.

I have been working on a music weighting script that performs two functions:

1) Determines a rating based on the song's skip and play count information, using proper statistics to account for the uncertainty from small sample sizes. It will also call attention to songs whose ratings are at odds with their actual play history so they can be reevaluated manually if desired (it will automatically do this, but it will take longer). Generating ratings is quite useful for large libraries where sorting music by hand is time consuming. Existing ratings and play and skip data are all taken into account during this process.

2) Determines a play order that gives preference to a song based on its play and skip history as well as the amount of time it's been since it's been played or skipped. It does so in a properly randomized fashion to prevent similarly rated songs from clumping together and to account for uncertainty in the data. This mode is compatible with custom playlists as it will simply pay attention the weighting for the songs contained in the list. It also adjusts the level of advantage popular songs have over unpopular ones to maintain a user defined level of songs successfully being played.

The program is written in Applescript and runs on Macs via iTunes. It adjusts the last played date of each song which allows iPods and iTunes to sort based on that simple parameter. The script can easily be set up to run automatically every few days. Note that almost all music players do not keep play and skip count information and would therefore be incompatible with my script, so it's pretty much Apple devices only. I've tested this script quite extensively, and if anyone is interested in giving it a try leave a comment and I'll write up documentation for it. Since proper docs are a lot of work I'll hold off until there is a request. If requested, I could make a version that only generated song ratings from the play / skip info. The working assumptions are different for a one time examination than a continuous process, but I've already worked out the algorithmic differences.

Theory:

The script is built around the binomial distribution defined by each song's play and skip count information. The distribution is used as the basis for a non-uniform random number generator to compute a weight, which is then multiplied by the amount of time since the song was last attempted (ie played or skipped). While the math for the cumulative binomial distribution is known, the inverse has not been solved explicitly, so a binary search is used to generate the random values.

To control the global success rate when playing songs, the weight generated by the non-uniform random number generator is raised to a power. Since all the weights are in the range from zero to one, the resulting value is also in that range, and values closer to one will decrease less in size. This gives, on average, and advantage to more popular songs. As an example:

The mean ratio for a one star song is 0.1, and the mean for a five star song is 0.9.

This means that normally a five star song will play .9/.1 = 9 times more frequently.

If the results are raised to a power of 1.5, one gets: (.9^1.5) / (.1^1.5) = 27 times more frequently.

The greater the exponent becomes, the larger the advantage higher rated songs will have. In theory, as a lot of data is collected on how popular each song is such a system shouldn't be required, but it can take quite some time for that to occur. It's important not to raise the exponent too high, or else lower rated songs and those simply lacking any information will very rarely be played. The user can set the desired range. The script waits until at least 50 songs have been played or skipped and then uses the ratio of success to guess what the exponent should be adjusted to. The last several runs are averaged together to prevent sharp changes in the output.

To determine the song's rating the binomial distribution formed by the song's play and skip count history is used to determine a rating that is greater than at least 5% of the distribution's total area (this threshold is user controllable). If a manual rating has been set, the script will slowly move the rating away from that value as the data shows it is unlikely. If no rating is provided, a neutral rating of 0.5 is assumed.

The rating doesn't actually effect the weighting algorithm, since the binomial distribution formed by the play and skip counts is used. To incorporate a user's manual rating, therefore, virtual plays and skips are added to the song's actual ones to change the distribution. The relative weight of the rating can be set, and a separate weight is used for computed ratings (iTunes shows hollow stars if a song's album has been rated but the song itself has no rating). The virtual counts are maintained separately so they can be discarded if the real data shows the user's rating to be flawed. If the user changes the rating to one incompatible with the existing data, the data is cleared and the song tested from scratch. The assumption is that the user has observed the data and disagrees with the conclusion for some reason (perhaps someone else had been using their iPod!).

One idea for improvement I'm considering is having the script generate album ratings based on data collected on that album's (or maybe even just artist's) songs. In theory this could allow songs to converge more quickly. It wouldn't count for as much as a manual user rating, but it doesn't require any extra effort. The best system would be for Genius to generate projected ratings, much as Netflix does, but Apple's lousy Genius paradigm is the subject for another post.