Article: Stereotype threat, gender and mathematics attainment: A conceptual replication of Stricker & Ward

What is stereotype threat?

Mathematics education researchers have long been concerned that mathematics is experienced differently by men and women. This concern is, in part, fueled by gender differences in post-compulsory participation rates in mathematical study and STEM careers. One mechanism which some believe contributes to these observed gender differences in participation is stereotype threat. This account suggests that members of negatively stereotyped groups underperform when that stereotype is salient, perhaps because stereotype-related thoughts place an extra burden on stereotyped individuals’ cognitive resources.

Abstract

Stereotype threat has been proposed as one cause of gender differences in post-compulsory mathematics participation. Danaher and Crandall argued, based on a study conducted by Stricker and Ward, that enquiring about a student’s gender after they had finished a test, rather than before, would reduce stereotype threat and therefore increase the attainment of women students. Making such a change, they argued, could lead to nearly 5000 more women receiving AP Calculus AB credit per year. We conducted a preregistered conceptual replication of Stricker and Ward’s study in the context of the UK Mathematics Trust’s Junior Mathematical Challenge, finding no evidence of this stereotype threat effect. We conclude that the ‘silver bullet’ intervention of relocating demographic questions on test answer sheets is unlikely to provide an effective solution to systemic gender inequalities in mathematics education.

Female participants received a higher mean score in the math competition. Contrary to the hypothesis, the "gender-first" female participant group received a non-significantly higher mean score than the "gender-last" equivalent, rather than having a significantly lower mean score as expected.

Participants’ mean scores, split by answer-sheet version and gender, are shown in Fig 2. As stated in our preregistration, these scores were subjected to a 2 (version) by 2 (gender) between-subjects Analysis of Variance (ANOVA). This revealed a significant main effect of gender, F(1,1165) = 8.410, p = .004, η2 = .007, which reflected that female participants had a higher mean score than male participants, 46.2 versus 42.7, d = 0.177. There was no significant main effect of version, F(1,1165) = 1.586, p = .208, η2 = .001, (means 45.7, 44.0, d = 0.091. Crucially, we did not find the hypothesized version-by-gender interaction effect, F(1, 1165) = 0.525, p = .469, η2 = .000. Indeed, contrary to the prediction of the stereotype threat account, female participants in the gender-first condition had slightly (but non-significantly) higher scores than those in the gender-last condition, 47.3 v 45.0, t(717) = 1.61, p = .108, d = 0.120.

To be consistent with Stricker and Ward, our primary preregistered analysis involved performing an ANOVA. However, we also ran a multilevel analysis to take account of between-school variation. We compared models with (i) random intercepts (where intercepts were able to vary across schools) and (ii) random intercepts and slopes (where both intercepts and slopes were able to vary across schools). Allowing intercepts to vary across schools yielded a significantly better fit (BIC = 10037.65) than a model where intercepts were identical across schools (BIC = 10298.66), χ2(2) = 268.1, p < .001. However, allowing slopes to vary across schools did not significantly improve the fit of the model that included version, gender and the version-by-gender interaction (both BICS = 10047.86), χ2(2) = 0.00, p = 1. In this model (in which gender was coded 0 for males, 1 for females; and version was coded 0 for gender-first and 1 for gender-last), neither version, b = -0.902, t(1161) = -0.551, p = .582, nor gender, b = -2.99, t(1161) = -1.863, p = .063, nor the version-by-gender interaction effect, b = -1.04, t(1161) = -0.499, p = .618, were significant predictors of participants’ scores (intercept b = 46.95, t(1161) = 8.96, p < .001). In sum, analyzing the data in this fashion again provided no evidence of the hypothesized version-by-gender interaction.

Finally, we conducted a preregistered Bayesian version of our main ANOVA. This required us to specify a model for the null hypothesis. As specified in our preregistration, we ran two analyses, with Cauchy prior widths of 0.2 and 0.5. Both analyses provided strong support for the model that only included gender as a predictor over the model which captured the predicted stereotype threat effect (i.e. the model which included gender, version and the version-by-gender interaction effect), BF01s = 8.123, 46.052 respectively.

One of the included schools was a high achieving all-female school, which has led to an increase of the female participant mean score. An exploratory analysis which excluded this school yielded a similar result to the preregistered (pre-study) analysis:

To explore whether our inclusion of a single-gender school in the sample effected the results (perhaps, for example, students at single-gender schools are not as affected by societal stereotypes as those at coeducational schools [cf. 24]), we conducted an exploratory analysis with the 329 participants from this school omitted. This resulted in an essentially identical pattern of results. In particular we again found no significant version-by-gender interaction effect, F(1,836) = 0.059, p = .809, η2 = .000.

To explore whether or not our decision to use the standard method of scoring the JMC affected the results, we conducted the primary ANOVA analysis again using (i) number of problems answered correctly and (ii) percentage accuracy (number of problems answered correctly as a percentage of problems attempted) as dependent variables. Our primary conclusions remained for both these dependent variables. Specifically, neither version-by-gender interaction effects with these two dependent variables was significant: number correct, F(1,1165) = 0.253, p = .615, ηp2 = .000; percentage accuracy, F(1,1165) = 0.326, p = .568, ηp2 = .000.

In sum, we found no evidence in favor of the hypothesis that female participants scored lower when they received the gender-first version of the answer sheet in any of our analyses, and a Bayesian analysis provided strong evidence against this hypothesis.

When researchers excluded the data from the all-female school, the researchers found a small male advantage that is consistent with the general population. Although the mean score was lower for female participants this time around, the researchers found no evidence of the hypothesized stereotype threat effect.

Might the small female advantage found in our sample, compared to the small male advantage found nationally, account for the lack of a stereotype effect in our data? Again, we doubt this. This difference was driven by the inclusion of a high achieving girls-only school in our sample. This school had the highest mean score of any which participated. When this school was excluded from our analysis, we found a small male advantage consistent with the overall picture, t(838) = 2.391, p = .017, d = 0.165. As noted above, our substantive conclusions remain if the analysis is conducted on this restricted sample (N = 840).

Although the study does not explicitly deny the existence of a stereotype threat, it rules out the effect of relocating demographic questions with respect to the gendered mathematics achievement gap, and also calls into question the robustness of the stereotype threat effect.