Society & surveillance
What researchers could infer from your Facebook Likes
A 2013 experiment showed how ordinary clicks could become sensitive predictions. Understanding its accuracy measures is essential to understanding the privacy problem.
A profile made from small choices
A Like looks like a small, voluntary disclosure: a musician, a television programme, a joke. The unsettling possibility is that a collection of such choices can answer questions the person never intended to answer. In 2013, Michal Kosinski, David Stillwell and Thore Graepel tested that possibility with 58,466 consenting American volunteers. They paired Facebook Likes with profile information and questionnaires, then trained statistical models to predict personal attributes.
The study tested predictions on people held out from model fitting. That matters: a system that merely memorizes its training examples has not demonstrated that its associations travel to someone new. The researchers reported substantial predictive information for some attributes, with much weaker performance for others.
What the headline percentages mean
For two-category attributes, the paper used the area under a receiver-operating-characteristic curve, or AUC. Its reported 0.88 for distinguishing male sexual-orientation categories is a ranking measure. It is not a claim that 88 percent of every population could be correctly labelled at any chosen decision threshold.
Imagine ranking one person from each of two categories. AUC asks how often their order is correct. An operational classifier must also decide where to draw a line, and the proportions of the two groups affect how many positive predictions will be wrong. A strong ranking score can therefore coexist with consequential mistakes. No percentage in the paper authorizes an inference about a particular reader.
A later test asked a different question
A 2015 study by Wu Youyou, Kosinski and Stillwell compared computer predictions with acquaintances' judgments of personality. Against participants' own questionnaire scores, the computer predictions correlated more strongly on average than the human judgments. This extended the research programme, but its target was agreement with a personality measure, not unlimited knowledge of a person.
A questionnaire is itself a measurement instrument. Predicting its answers is different from understanding someone's reasons, and a volunteer sample is different from the entire population. Both papers should be read as studies of bounded prediction tasks. Neither is an investigation of Cambridge Analytica's later operations.
The inference is the disclosure
The practical lesson is about the gap between collecting a datum and using it. A person might willingly share a music preference while rejecting its use in a decision about work, insurance or political persuasion. Those are separate choices, even if the second system never asks a sensitive question directly.
There is also a double failure to consider. A correct inference can expose something private; an incorrect inference can attach a damaging label. Evaluating such systems therefore requires more than a single accuracy score. The relevant questions include who receives the prediction, how uncertainty is shown, whether the person can challenge it, and what consequences follow.
Sources and further reading
- Private traits and attributes are predictable from digital records of human behavior ↗
PNAS 110(15), 5802–5805; Figure 1, Results, Figures 2–4
- Youyou, Kosinski and Stillwell, Computer-based personality judgments (2015) ↗
Abstract; PNAS 112(4), 1036–1040