The Limits of Expert Prediction — Reading
Passage
A In 1984 the psychologist Philip Tetlock began an experiment that would run for twenty years. He recruited nearly three hundred people whose profession involved commenting on political and economic trends — academics, journalists, government advisers — and asked them to assign probabilities to specific future events. Would a particular government still be in office in five years? Would a given economy grow, shrink or stagnate? By the time the study closed he had gathered more than eighty thousand forecasts, each one precise enough to be scored against what actually happened. B The headline result has been quoted, and misquoted, ever since. On average, the experts performed slightly worse than simple statistical rules extrapolating from past data, and on some questions barely better than chance. What is usually left out is that the average concealed an enormous spread. A minority forecast considerably better than the rest, consistently and across domains. The interesting question was never whether expertise helps, but why it helped some people and not others. C Tetlock borrowed a distinction from Isaiah Berlin, who had divided thinkers into hedgehogs, who know one big thing, and foxes, who know many small things. The hedgehogs in his sample organised everything around a single governing idea and extended it confidently into unfamiliar territory. The foxes assembled explanations from several partial frameworks, held them loosely, and were quicker to abandon one that was failing. It was the foxes who scored better. Strikingly, the hedgehogs were the ones broadcasters preferred: a commentator who reduces a complex situation to one clear principle makes better television than one who begins by listing what is uncertain. D Confidence turned out to be the clearest warning sign. The forecasters who were most certain were, on average, least accurate, and the effect grew stronger the more famous the forecaster. Tetlock also documented how poorly the confident handled being wrong. Rather than adjusting, they reached for defences — the prediction was nearly right, or would have been right had some unforeseeable event not intervened, or concerned a timeframe that had not yet elapsed. These moves preserve the underlying theory intact, which is precisely why they prevent learning from failure. E Not everyone accepts the pessimistic reading. Some critics argue the questions were unrepresentative, chosen because they were scoreable rather than because they were the questions expertise is actually for. An economist may have no idea which party will win an election while understanding perfectly well how a currency peg fails. Others note that in domains with fast, unambiguous feedback — weather forecasting, competitive bridge, anaesthesiology — expert judgement is demonstrably excellent. On this view the finding is not about expertise but about environments: skill develops where the world corrects you promptly, and fails to develop where it does not. F Tetlock's later work took that criticism seriously. A forecasting tournament run for the American intelligence community from 2011 identified individuals who substantially outperformed the field, including analysts with access to classified material. These 'superforecasters' shared habits rather than credentials: they broke large questions into smaller ones, began from base rates before adjusting for specifics, updated in small increments as evidence arrived, and worked in teams that argued. Training in these habits measurably improved accuracy, which suggests the trait is a practice rather than a gift. G The practical conclusion is narrower than either the sceptics or the enthusiasts would like. Expertise is not worthless, and confident expertise is not automatically suspect. But the credential that qualifies someone to explain why something happened does not, on its own, qualify them to say what will happen next — and the manner in which a forecast is delivered carries almost no information about whether it is any good. H Prediction markets, in which participants trade contracts whose payout depends on the outcome of a real-world event, are sometimes proposed as a structural fix for exactly this problem, on the theory that requiring forecasters to back their claims with money — rather than merely their reputation — filters out the confident but poorly calibrated commentary that dominates televised punditry. Aggregated market prices have in several documented cases outperformed both individual expert predictions and opinion polls, particularly for events with a large number of informed participants and a clear, unambiguous resolution criterion. The method has real limits, however, that prevent it from simply replacing expert judgement altogether: markets require enough informed participants trading real stakes to function well, they perform poorly on questions with ambiguous or long-delayed resolution, and they can be distorted by well-funded participants trying to manipulate the price rather than genuinely forecast the outcome, particularly in markets tied to politically contested events. What the method demonstrates most convincingly is not that markets are infallible, but that forcing a forecaster's confidence to be backed by a real cost for being wrong produces better-calibrated predictions than a media environment that rewards confident delivery regardless of a forecaster's actual track record.