Social psychology’s most influential theory says mixing different groups reduces prejudice. A re-analysis of 31 studies and 33,000 people found the causal evidence for that is largely a statistical artifact
The contact hypothesis has shaped more real-world policy than almost any other idea in social science. It is the theoretical foundation behind school integration programs, mixed-housing developments, workplace diversity initiatives, and international exchange schemes. First proposed by Gordon Allport in 1954 and subsequently supported by hundreds of studies involving more than a quarter of a million participants, the idea is simple and appealing: when people from different groups meet, interact, and get to know each other, they come away with less prejudice than they had before.
A new study published in Communications Psychology has examined the longitudinal evidence for that claim more carefully than it has ever been examined before, and found that the conclusion depends almost entirely on which statistical model you use to analyze the data. When researchers applied the methods that social psychology has traditionally used to test whether contact causes attitude change, the effect appeared. When they applied more rigorous methods that better account for the factors that determine who interacts with whom in the first place, the effect largely disappeared.
The paper does not claim that contact never matters. It claims that the causal evidence for it, the evidence that contact itself is what produces attitude change rather than pre-existing similarities between people who choose to interact, is considerably weaker than seven decades of research has suggested.
What the contact hypothesis actually claims
Allport’s original formulation was careful. He argued that contact between groups would reduce prejudice under specific conditions: equal status between the groups, common goals, institutional support, and cooperative rather than competitive interaction. The research that followed, particularly the landmark meta-analysis by Pettigrew and Tropp in 2006 covering 515 studies and 250,493 participants, concluded that the effect held even without those optimal conditions, and that contact between groups was robustly associated with reduced prejudice across a wide range of settings.
But association is not causation. The central difficulty in contact research has always been the same: people are not randomly assigned to interact with members of other groups. The individuals who end up living in integrated neighborhoods, attending integrated schools, or working in diverse offices are different in systematic ways from those who do not. They may already hold more favorable attitudes. They may have more education, higher socioeconomic status, or personality traits that predict both more intergroup contact and more openness to outgroups. If these pre-existing differences are not fully controlled for in the analysis, the observed association between contact and reduced prejudice may reflect who was likely to have contact, not what contact actually does.
Cross-sectional studies, which measure contact and attitudes at one point in time, cannot disentangle this problem at all. Longitudinal studies, which follow the same people over time, can in principle do better, because they can observe whether changes in contact are followed by changes in attitude in the same individuals. But even longitudinal studies depend on statistical models to isolate causal effects, and the model choice turns out to matter enormously.
The model that social psychology has been using
The standard statistical tool for analyzing longitudinal contact data is the cross-lagged panel model, known as the CLPM. In a CLPM, researchers measure both contact and attitudes at multiple time points and test whether earlier contact predicts later attitude change, controlling for earlier attitudes. If higher contact at time 1 predicts more positive attitudes at time 2 even after controlling for attitudes at time 1, the CLPM interprets this as evidence that contact caused the attitude change.
The CLPM has been used in the vast majority of longitudinal contact studies, and when the researchers applied it to their combined dataset of 31 longitudinal datasets from 21 published studies, it produced the expected result: prior contact predicted more positive later attitudes with statistical significance. The contact hypothesis appeared confirmed.
But the CLPM has a known statistical problem. It assumes that all participants start from the same baseline and that any differences between individuals are fully captured by the variables measured at the first time point. In practice, people who have more intergroup contact over time tend to differ from those who have less contact in ways that are stable across the entire observation period and that the first-wave measurement cannot fully capture. If these stable between-person differences are not modeled separately, they contaminate the apparent longitudinal effect of contact on attitudes.
What happens when the models improve
Two more recently developed statistical approaches address this problem directly. The full-forward cross-lagged panel model and the random-intercept cross-lagged panel model both separate the within-person dynamics that the CLPM is trying to capture from the stable between-person differences that it conflates with them. By modeling each person’s stable baseline separately, these methods can estimate more precisely whether changes in contact within the same individual over time are followed by changes in that individual’s attitudes.
When the researchers re-analyzed the same 31 datasets using these improved models, the longitudinal contact effect shrank substantially and, across most analyses, became statistically non-significant. The same data that showed a significant causal effect under the CLPM showed no significant longitudinal effect under the models better equipped to identify genuine within-person change.
What remained was a large and consistent between-person correlation: individuals who reported more contact over the observation period also tended to hold more positive outgroup attitudes over that same period. But a between-person correlation, no matter how robust, cannot be interpreted as evidence that contact caused the attitudes. It could equally reflect that people who already hold more positive attitudes seek out more contact, or that some third factor, education, personality, socioeconomic position, neighborhood self-selection, produces both more contact and more positive attitudes simultaneously.
“The same data that appeared to confirm the causal story under traditional methods told a much more cautious story under methods that better controlled for unobserved differences between people,” the authors write. “That distinction matters enormously for how the field should interpret its evidence and design its interventions.”
What this does and does not mean
The researchers are explicit about the limits of what their re-analysis establishes. It does not prove that contact has no effect on attitudes. It establishes that the longitudinal evidence for a causal effect is much weaker than previously believed, because that evidence was produced by methods that could not adequately distinguish cause from correlation.
The possibility that genuinely high-quality contact, under the specific conditions Allport originally specified, does produce attitude change in the individuals experiencing it remains open. What the study challenges is the empirical foundation for the strong version of the contact hypothesis, the claim that intergroup contact reliably produces attitude change across the broad range of naturalistic settings studied in longitudinal survey research.
Several important distinctions survive the re-analysis. The between-person correlation between contact and positive attitudes is real and large. More contact is associated with better attitudes, consistently across datasets. The dispute is about whether that association reflects a causal process that contact interventions could reliably reproduce, or whether it primarily reflects pre-existing differences between the kinds of people who have more or less intergroup contact in everyday life.
The practical implications of that distinction are significant. If the contact effect is genuinely causal, then designing programs that increase intergroup contact should produce measurable attitude change. If the association is primarily a reflection of pre-existing selection, then more contact may not change attitudes much for people who would not have sought it out on their own. The interventions that look most effective in observational data may be working for other reasons, or may be working only in the specific populations and conditions from which the data came.
The broader crisis in social psychology
The contact hypothesis sits within a broader context of replication difficulties that have affected social psychology over the past decade. Many effects that appeared robust in the original literature, from priming effects to ego depletion to stereotype threat, have been found to be smaller than originally reported, context-dependent in ways that limit generalization, or dependent on methodological choices that produced inflated effect sizes.
The pattern the contact study documents is not a simple failure to replicate a specific finding. It is a demonstration that the most widely used analytical tool in an entire subfield, longitudinal contact research, systematically overstates causal effects by conflating stable individual differences with the within-person dynamics it was designed to capture. Every published longitudinal study that used the CLPM to support the contact hypothesis is subject to the same concern.
This does not mean those studies were wrong in identifying an association. It means the field may have been too quick to interpret that association as causal, and to build policy recommendations on an inference the data were not able to support.
The researchers frame their conclusion as a call for better methods rather than an abandonment of contact theory. Future research using randomized designs, where participants are assigned to contact conditions rather than self-selecting into them, would provide the kind of causal evidence that longitudinal observational data cannot. Experimental contact studies using random assignment have produced some evidence of attitude change, and those designs are more credible than the observational longitudinal data the current re-analysis examined.
What the contact hypothesis ultimately rests on, after seventy years of research and more than thirty-three thousand participants in longitudinal studies, is a between-person correlation that is real but causally ambiguous, and a causal claim that the most rigorous available methods cannot confirm.
Source
Friehs, M.T., Schäfer, S.J., Wüst, K. et al. “A comparative multi-method re-analysis of the longitudinal evidence for causal intergroup contact effects on attitudes.” Communications Psychology, 4, 107 (2026).
DOI: 10.1038/s44271-026-00495-8