Showing posts with label bonus question. Show all posts
Showing posts with label bonus question. Show all posts

Tuesday, 8 July 2014

Bonus question
Was Series 9 of Only Connect the toughest yet?

Only Connect had its last hurrah on BBC 4 last night with the end of series 9 heralding its long-anticipated move to BBC 2 in the autumn. Much has been made of the supposed threat this might pose to the show's famed difficulty, with many a doubter worried about an inevitable 'dumbing down' as it tries to find its feet in more populist waters. Conversely, the latest series - featuring a brand new question editor - has been cited as the 'toughest ever', with some going as far as to say the questions have been "impossible", "unfair", and "a masterpiece in obscurity". Now that the series is complete what can we take from the scorecards the latest inductees to the Only Connect Cadet Force have turned in?

Our first exhibit is the simplest: average total score (of both teams) per episode for each series, contrasted with the series 1-8 average.
Average (combined) per-episode scores of Only Connect teams, series 1-9.

It probably won't come as too much of a surprise that, with an average combined score per game of just under 34 points, series 9 is indeed the lowest scoring yet. What's slightly interesting, however, is that this puts it just a couple of points lower than the previous record holder: series 2's teams averaged 36 points per game between them. Nevertheless, series 9 is rather out on its own and a standard statistical test suggests that this fluctuation isn't just down to chance, so something's going on (and for those of you who yearn for p-values, it's 0.01).

It's fairly clear, then, that scoring this series was abnormally low, but it gets a little more interesting if we look at round-by-round scoring to see exactly where these points have been lost. The next graphs split average scores per game into individual rounds (you may want to click it to make it slightly bigger, or even right-click and open in a new tab).

Average (combined) per-episode round-by-round scores of Only Connect teams, series 1-9.


Series 9 doesn't stand out quite so much on these plots, as there is understandably rather more natural variation in scores when we look at things in more detail. Nevertheless, series 9 saw the lowest average scores in sequences and on the walls, and was also well below average for connections and missing vowels to boot. Compared with the Series 1-8 average, Series 9 episodes saw around 1 fewer point scored in rounds 1 and 2, a 2.5 point drop on the walls and another 2 points or so in missing vowels. While this points to a general across-the-board fall in scoring, the wall scores are responsible for slightly more than their fair share of the drop-off.

Let's take a more detailed look at those wall scores. The following compares the distribution of scores on individual walls for series 1 to 8 with those of series 9: the height of a bar indicates what proportion of walls were solved for that particular score. It's here where a big part of series 9's lower scores is hiding.

Particularly striking are the bars at the far end. Out of series 9's 26 walls just 4 - under 1 in 6 - were solved for the maximum of 10 points. In contrast, across the 224 walls in series 1 to 8 a whopping 78 - over 1 in 3 - were maxed. While comparing the relatively small series 9 dataset with series 1-8 leaves a lot of room for statistical noise to creep in, it's still a fairly remarkable change in scoring, and again there is some reasonable statistical evidence this isn't just down to chance (p = 0.02, stats-fans).

Go hard or go home?

So where does this little tour of Only Connect numbers leave us? One thing that's inescapable is that the scoring in series 9 of TV's toughest quiz was not merely the lowest yet, but also so low that the odds of this being down to random chance alone are low enough to interest a statistician (and we're very interesting people). In other words, there's reasonable evidence that something is underpinning the drop. What exactly is, however, debatable. Harder questions, perhaps? Or could it simply be this series' contestants being a touch sub-par?

The data, alas, cannot distinguish these two possibilities (much like the never-ending debate over whether exams are getting easier or kids are getting smarter). Personally, though, I'm largely in the "ouch, that was a bit tricky" camp. Like every one that has preceded it, this series featured contestants with some serious quizzing pedigrees along with some entirely new faces, and I certainly don't think any of them were shown to be chumps (I'd welcome anyone who says otherwise to apply!). On the other hand, the questions have seemed a touch stiffer from the comfort of my sofa, and I think the change in question editor has shown through in a slight shift in the styles of puzzles we've been faced with. How this will develop for the show's first (of hopefully many) series on BBC 2 remains to be seen, but for now at least I think fears of Only Connect going soft can be put to one side.

Although really, if there's one show that can afford to dumb down, it's this one.

Monday, 7 July 2014

Bonus Question
Who's going to win Only Connect?

Two teams, one trophy.
Doesn't time fly? After 12 weeks of mind-bending questions, Series 9 of Only Connect is about to reach its climax. Tonight we'll see the Europhiles take on the Relatives as they vie to stop me introducing myself as "reigning Only Connect champion". (Which inexplicably doesn't seem to work too well here in Canada.) What's more, for the last show on BBC 4 it will be an Only Connect first, with the final a rematch of a first round encounter which saw the Europhiles sneak home 17-14. This suggests the Relatives are tonight's underdogs, but I thought I'd look through the archives to see whether finalists' form across a series can tell us much about who's more likely to take home the trophy. (And if you're new to the blog, you may find my analysis here using the same data of how important missing vowels is interesting.)

For each of the eight Only Connect finals to date I compared the round-by-round scoring of the eventual champions with that of the runners-up across their respective runs to the final. For each team I looked at the following six statistics (averaged across their games): total points scored, winning margin in those games, points scored in each of the first three rounds (connections, sequences, and walls), and points percentage in missing vowels.

These are hopefully self-explanatory except for the missing vowels 'percentage', which is a stat I concocted for comparing series champions in a previous post. Rather than looking at total points scored in missing vowels, which can vary in length from show to show, I instead compare how many points a team scored with the combined total of their and their opponents' score in the round. For example, if a team scores seven points in missing vowels while their opponents score three, their missing vowels percentage would be seven out of ten, or 70%. Not ideal, of course, but as most missing vowel clues are solved it's hopefully slightly more reliable than looking at raw score for that round.

For simplicity I decided to consider, for each of these metrics, whether or not the team that went on to become series champions had a better or worse record across the series than the eventual runners-up. The results of this can be seen in the figure below, where green indicates the champions had a better record, red a worse one, and orange that they were all square. (If you have trouble with red and green I've made a greyscale one, and if you want to see the full numbers you can check them out in an ugly table here.)


The first thing to note is that in every series so far the eventual champions have scored more points on their way to the final than the runners-up. Champions also tend to have better winning margins, with this being the case for seven of the eight. The round-by-round data, meanwhile, are rather less informative. Sequences are arguably the best indicator of potential champions, with six out of the eight champions demonstrating a better record on that front (and on the two other occasions things were very close), but beyond that you're usually looking at something of a coin toss.

How the Series 9 finalists compare.
So what of the latest pretenders to the Only Connect throne? Unique in having a head-to-head record we can already say that the Europhiles are 1-0 up, but if we look at their average records this series things start to turn the other way.

The Relatives have the edge in four of the six metrics I've considered, including the apparently all-important 'average points' statistic, having averaged an additional 3.5 points per show. They only trail in winning margin and missing vowels, with the Europhiles' record in the latter particularly impressive. However, while the Europhiles have only conceded a grand total of six points to opponents in missing vowels, five of these came in their early encounter with the Relatives. If things are close going into the final round be prepared for some fireworks.

Now before I get a bunch of angry comments telling me to give back my PhD I should acknowledge that all this analysis is a little tongue-in-cheek given the size of the dataset I've worked with and the numerous assumptions I've had to make. The biggest question mark hangs over whether it's at all legitimate to compare team's head-to-head records, as these will be at least partly determined not only by which opponents they happened to meet along the way, but also by their route through the tournament. With later rounds supposedly featuring harder questions, for example, the Europhiles have presumably faced tougher sets on average given the Relatives' extra game during the theoretically easier first round. Then again, the Relatives pulled off an impressive 18-15 win in their semi-final while the Europhiles came home 11-7, but it's impossible to say how much of this is down to a better team, and how much is down to easier questions or just plain old luck of the draw.

Nevertheless, the stats alone suggest that while the Relatives may have stumbled in their first meeting they're now firmly the favourites on paper. My money, however, is on the Europhiles, who have impressed me a touch more over the series (and because you can rely on a statistician to hedge his bets). Either way we should be in for a cracking contest, and once you've seen it come back tomorrow where I'll be looking at the real Only Connect talking point: "is it just me, or was this series really, really hard?".

Friday, 13 June 2014

Bonus Question
World Cup Quiz!

Yes, it's the (men's) World Cup, and because I've been far too busy being far too excited about football this week I've not had time to do any pub quizzes. But fear not: instead of the usual quiz questions we got wrong, I've put together a quick World Cup Quiz. A mere four rounds (including pictures, naturally) I've tried to ask a mixture of things that are either interesting, or things that are boring but nevertheless essential trivia knowledge for all those World Cup questions quizmasters across the country will be asking over the next month. I'll have doubtless missed a bunch of obvious material so if you have any favourite football factoids do please let me know, either here or via Twitter @statacake!

Warming up
1) Which country both hosted and won the first World Cup in 1930?
2) Even non-fans know that Brazil are very good at winning the World Cup, but (prior to the 2014 competition) how many times have they won it?
3) Before this tournament, which player holds the record for most goals scored in World Cup Finals matches? Germany's Miroslav Klose is just one goal behind though, and has the chance to equal or even overhaul that record this year.
4) What was the final score in the only World Cup Final that really counts (1966)?
5) What is the official name of the current World Cup trophy?

The answers


Some stretches
1) The 2010 World Cup ball was criticized by some (especially goalkeepers) for being 'too round'. Meaning 'celebrate' in Zulu, what was that ball called?
2) England have participated in three penalty shootouts at the World Cup Finals and, you guessed it, have lost all of them. In the process seven England players have suffered the ignominy of missing a World Cup shootout penalty. Name five of them.
3) Complete this quote from Luis Suárez following his team's quarter-final victory in 2010: "The _____ now belongs to me".
4) The name 'Fuleco', chosen for the Armadillo mascot of this year's tournament, is a portmanteau of two Portugese words. What do these two words mean in English?
5) What was notable about the sequence of World Cup champions from 1970 to 1994, leading some pundits to suggest that England were 'destined' to win in 1998?

The answers


Disciplinary matters
1) To date, five players have been sent off in a World Cup final, but which Frenchman was the most recent to receive this particular 'honour' for headbutting an opponent?
2) ...and who did he headbutt?
3) A mere three England players have been sent off during a World Cup Finals match. Name two of them.
4) A match in the 2006 finals earned the nickname 'The Battle of Nuremberg' after it saw a record four red cards and 16 yellow cards dished out. Name either of the teams involved (and no, Germany wasn't one of them).
5) Which year's World Cup Finals saw the first introduction of red and yellow cards in professional football? This is an important trivium to keep in mind for when a quizmaster asks "Who/when was the first player shown a red card in a World Cup Finals match?" instead of "Who/when was the first player sent off in a World Cup Finals match?"

The answers


Picture Round!
Finally, some pictures. Here are five World Cup mascots, can you match them up to their respective host countries? Any helpful wording has been expertly removed.


The answers

Tuesday, 3 June 2014

Bonus Question
Does the missing vowels round still not matter?

Still my proudest accomplishment.
Back in October I put up a post addressing the regular complaint that Only Connect's missing vowels round is 'overpowered' using data from the first seven series of the show. A mere six months after it concluded I thought I'd stick up a quick summary of how adding Series 8's data affected things. The answer? Not much. Still, read on for the latest Only Connect missing vowels stats (until Series 9 finishes in a couple of months and I end up doing it all again).

Turnabout's fair play

For me, the headline statistic when it comes to the importance of missing vowels is turnarounds: how often does a team come from behind after the walls to snatch victory from the jaws of defeat? Across the first seven regular series of Only Connect - comprising 99 episodes - 14 teams achieved this feat. Series 8 added three to this total, taking us to 17 turnarounds in 112 shows, or about twice every 13 episodes. There have still only been four occasions where teams were tied after the walls. (For the extra-curious, all three of the new additions took place in the first four episodes of Series 8, with the Lasletts and the Bakers overturning single-point deficits to defeat the Pilots and Press Gang, respectively, while my team the Board Gamers came back from three behind to overtake the Globetrotters.)

Series 4's Radio Addicts, one of just three teams
to turn around a post-wall deficit of four points.
Unsurprisingly, this hasn't changed another important stat: the biggest missing vowels comeback remains a relatively meagre four points, with this having been accomplished just three times in the show's history. My team was the third to turn things around from three points behind, while four teams have managed a 2-point turnaround. Over half of all the missing vowels comebacks have been from just 1 point behind. If missing vowels is an elaborate scheme to help quick-fingered teams win, it's not a very successful one.

Swing when you're winning

Points swings per round: how much the
team that wins a round will win it by.
A new statistic I thought I'd look at this time is points 'swings' per round: when a team wins a round, how many points they win that round by. This gives a sense of 'volatility': how many points the better team might claw back, or extend their lead by, in each of the four rounds. Doing this we get the table on the right where (for example) we see that on an 'average' episode one team will score about 2.6 more points than the other in the connections round, 3.4 on sequences, 2.4 on the walls, and 4.2 on missing vowels. (For the sake of completeness, the relevant standard deviations tell a near-identical story.)

Of obvious note is the missing vowels swing: on average the team which wins that round will win it by more than any other, albeit by less than one point over sequences. However, the fact that missing vowels so seldom makes a difference to the final outcome of a match suggests that while on average one team will score about 4 points more than their opponents in this round, they will usually either be so far ahead - or behind - that it doesn't matter. Missing vowels has the potential to cause big upsets, but only if one team is exceptionally good at them while simultaneously being rubbish at the rest of the show. Moreover, with so little to choose between the four rounds in absolute terms, the same could be really be said of any of them. (Plus, ultimately, like-for-like comparisons such as this are tricky given the fundamentally different nature of scoring between the rounds.)

Summing up

Finally, an update to the average points per round stats. As the table below shows, Series 8 doesn't stand out at all, with the overall averages largely unchanged (desipte my team's best efforts to drag the walls average down...).

Average points scored (by both teams) per round on Only Connect Series 1-8. (Click for big.)
That's your lot. In short: the missing vowels round is still nowhere near as overpowered as plenty of people seem to think: over 80% of the time it makes no difference to a show's outcome, and even then it will likely only see a one point deficit overturned.

Wednesday, 11 September 2013

Bonus Question
Quizzing Grand Slams?

A curious property of TV quizzes is that there is often little correlation between effort and reward. Some shows offer huge prizes in return for particularly exceptional quizzing, while others reward only mild skill (or even plain old luck) with similarly enormous paychecks. The most peculiar extreme, however, features those shows that are the hardest by far to win, but offer virtually no prize whatsoever. Compare Mastermind, one of the toughest gigs in town, with Deal or No Deal, where contestants are asked (the same) yes or no question a few times. On the latter, and on a daily basis, players could win up to £250,000, on the former one contestant a year gets a fruit bowl.

Clearly then, people aren't applying for Mastermind for the money. Instead there's some intangible prestige associated with claiming this particular televisual title that keeps the contestants rolling up. Tell some hardcore quizzers you won a few quid off the Banker and they might feign mild interest, tell them you've won Mastermind and, well, you won't need to tell them because they'll have already tried to recruit you to their quiz league.

During a recent pub quiz someone suggested that Mastermind could be regarded as a 'quizzing Grand Slam'; one of those few tournaments where the title means more than any prize money that might come with it. (And while I'm well aware of how this sounds just a little bit silly in the cold light of day, it honestly made perfect sense after a few pints.) What followed was an inevitable debate over which other quiz shows could be considered Grand Slams, and whether we could pair them up with their real-life equivalents in tennis (other sports are available, but it was Wimbledon season). We decided that to qualify a show had to have no prize money, be currently televised (thus ruling out Fifteen to One, an otherwise obvious contender), and be 'bloody hard'. Your mileage may vary, but what follows is thus the entirely subjective view of a group of moderately squiffy quiz nerds.

"You should've chosen 'being a crybaby'
as your specialist subject!!!"
Wimbledon

The British classic can only be matched by one show: Mastermind. While not the oldest on the list, it is arguably the 'daddy' of TV quizzes with the greatest recognition even among those strange members of society who don't obsess over trivia. Sure, Andy Murray won the US Open last year, but it was his subsequent victory at Wimbledon that got the most attention. Win Mastermind and chances are even your hairdresser will be impressed.

The Australian Open

The youngest of the tennis grand slams, we thought University Challenge, with its inherently youthful flavour, was the best match. Another unashamedly tough quiz, it's possibly the hardest to win given the relatively strict eligibility requirement of being a student. While the Open University has seen some 'back door' routes into this title in the past, their lack of presence in the competition in recent years suggests that even that path may be closed for the time being.

The US Open

While the oldest of the lot, the radio-only presence of Brain of Britain means it doesn't quite have the Wimbledon-esque significance of Mastermind in the public consciousness. Nevertheless, there's no denying the difficulty and importance of claiming this particular accolade in the quizzing world. I'll admit the US Open analogy is more process of elimination than perfect match, but if you really want a tenuous link then let's say Brain of Britain's somewhat idiosyncratic question structure reflects the US Open's position as the only grand slam with final set tiebreaks. Uncanny.

The French Open
Rafael Nadal's skills at missing vowels
are a closely-guarded secret

With the French Open's clay courts setting it apart from the other three tennis Grand Slams, we felt the lateral thinking element of Only Connect made it a fitting final entry to the list. While perhaps too young to be a cast-iron consideration just yet, there is sadly little competition for this fourth and final spot. Whether it can stand up to the test of time remains to be seen.

Final thoughts

I don't doubt that many would argue with the above but I can at least claim not to be the only person to consider these shows as a quizzing 'big four'. Former Mastermind champion (not to mention Brain of Britain and Only Connect runner up) Dave Clark identified appearing on these four shows as one of '30 quiz experiences to try before you die' (which is a good read if you haven't already seen it), and various googling can throw up other similarly 'qualified' individuals espousing the virtues of every member of the list.

In any case they all certainly tick the boxes of having no prize money and being bloody hard. Indeed, with regards to the latter no-one has managed to win all four. To date Ian Bayley has come the closest, winning Brain of Britain in 2010, Mastermind in 2011, and is one third of Only Connect's all-conquering Crossworders. He also carries the rare distinction of having appeared on University Challenge twice, for two different institutions, but despite this is unable to complete the set. Clearly, if you want to achieve this holy quaternity of quiz shows, you need to start young.

Thursday, 2 May 2013

Bonus Question
Are Manchester really the best at University Challenge?

Monday saw the conclusion of the 2012-13 series of University Challenge where, much to my annoyance, Manchester avenged their quarter-final defeat to University College London to take home the trophy. It's the first time the title has been retained since Magdalen College, Oxford secured back-to-back victories in 1997 and 1998 (an institution Manchester now also join on four victories in the prestigious quiz).

Teams from Cottonopolis have become a familiar sight on Monday night BBC2 TV schedules of late: in the last eight years they've lifted the trophy four times and only failed to make the semi-finals once (when their team didn't make the cut to appear on the programme). There seems little questioning their status as the most consistent institution over the years, but do the data back this up? While we're at it, can we identify the Best Team Ever? It's time for a statistical adventure.

First up: what data? University Challenge first aired in 1963, running for 25 years before being taken off air in 1987. Picked up again eight years later, it has been a fixture of television schedules ever since, and with full round-by-round scores available from 1995 onwards it's this 'Paxman era' where we'll be focusing our attention.

Now we have some data, the next question is how to compare teams both within and across series. At a basic level things are straightforward: it seems fair to assume that the series champions were the best team in that series, while the runners-up were - by definition - second best. But how do you compare the two losing semi-finalists, or compare this year's winners to the champions from 1995?

There are, of course, numerous ways we could derive metrics to compare teams (indeed, there are plenty of established methods in existence) but I wanted to build my own, as-simple-as-possible, model based on three 'intuitive' principles:

1) Progressing further in the competition is better
2) Losing by a small margin is better than losing by a large margin
3) Losing to a team that goes on to do well is better than losing to a team that goes on to do badly

The first of these is made straightforward by the (relatively) consistent tournament structure on the show. Since 1995 every series has featured five rounds, so I decided to assign every team a Baseline score from 1 to 6 based on their stage of elimination from the show: 1 point for losers in the first round (or highest scoring losers who lose their playoff match), 2 if you went out in the second round, and so on up to 5 points for the losing finalists and 6 for the series champions. This measure is the first element of comparing teams within a series: a higher Baseline score means a better performance. The problem is how to separate teams who were eliminated at the same stage of competition, which is where we try and incorporate the second and third principles.

How far a team progressed in a series is one half of how good they are: the other, of course, is how they fared against - and the quality of - the opponents they met along the way. For opponent strength we have a ready-made statistic in their Baseline score, and we can use the scoreline from each of their games to see how well they did. A typical approach in tournaments is to look at the margin of victory or defeat - the 'spread' - but I decided instead to look at the proportion of the total points scored in a game that were picked up by either team. This means that the effect varying question difficulty across rounds (or even series) is moderated, and also gives us a handy metric of 'performance' in a game in the form of a percentage: if a team lose 150-50 then they picked up 25% of the points in that game, while if they were pipped 155-150 it would be almost 50%.

By multiplying the percentage of points scored in a game by the opponent's Baseline score, we get a measure of performance which I've imaginatively called Performance score. For example, suppose you lost in the first round 150-50 to a team who went on to win the series. Your opponent's Baseline score would be 6, while your points percentage for that game is 25%. Combining these gives you a Performance score of 25% x 6 = 1.5. Your opponents, meanwhile, bagged 75% of the points available, but as first round losers your Baseline score is just 1. They therefore pick up a Performance score of 0.75. It might seem a bit odd that you get more points for losing than they do for winning, but remember that this measure is only used to compare teams who were eliminated at the same stage of the competition, so this comparison doesn't really mean anything.

From here, we can calculate every team's average Performance score across all of their games, giving a measure of the strength of their opponents and how well they fared against them. We can then use this metric as a tie-breaker to separate teams who have the same number of wins. For example, if we apply this strategy for the current series, we find that of the two losing semi-finalists (New College, Oxford, and Bangor) Bangor would snatch third place. (Admittedly, I was a little surprised by this as New College seemed the much stronger team, but a quick look at the results for the series suggests that this isn't reflected in the scores. For example, Bangor defeated King's College, Cambridge, far more convincingly than New College did.)

In the same way we can also compare the 19 Paxman-era champions to see which team were the most dominant in their series. It will come as little surprise to regular viewers that the 2009 Corpus Christi, Oxford team (aka Corpus Christi Trimble) would have topped this particular list, but as they were disqualified for fielding an ineligible player we instead find the 1998 Magdalen, Oxford squad come out on top. This team were a little before my time, but a poke through that series suggests that the scoring algorithm is doing a reasonable job: their quarter- and semi-finals were Trimble-like demolitions before a relatively narrow victory against Birkbeck to lift the trophy. (Coincidentally, Magdalen also take second in the overall standings, with their 2011 team posting similarly strong statistics.)

What of our original question, though? Which institution has been the most successful at University Challenge in the last 19 years? For this I assigned every team a rank within their series (first based on how far they got in the competition then using average Performance score above to break ties). From here there are then two ways to identify the 'best' institution: their average rank or their total rank. If we go with the former then, predictably, it's a team with only one appearance who top the list: London Metropolitan may have only made it onto the show once, but their third place in the 2004 series gives them a hard-to-beat average. Really, though, as getting through the show's audition exam is itself an achievement, it's total rank that represents a truly consistent institution, and on this metric it's Manchester who take the crown. The top of this list is, however, dominated by teams with multiple appearances: with a whopping 15 appearances Durham are second despite never winning a series, while Magdalen, Oxford are down in seventh with 'only' 9 appearances.

So there you have it, unequivocal proof that Manchester are doing something right, although if you're not totally convinced by my methods I wouldn't blame you. Merely while writing this up I spotted at least half a dozen holes one could pick in my metric, and I fully anticipate being alerted to some better, more established approach. Still, it's hard to deny its simplicity, and in any case the most important thing is who comes out on top. I don't think there are many systems that would suggest anything other than what mine has here: Manchester will once again be the team to beat next year, and Corpus Christi wuz robbed.