The original paper, published in Infant Behavior & Development:

https://www.math.kth.se/matstat/gru/5b1501/F/sex.pdf

Infants were presented with a face and a mobile separately, in a randomized order. (See Fig. 1). Testing was carried out at the mother’s bedside or in the neonatal nursery, at the Rosie Hospital, the choice of location depending on which was quietest. Overhead lighting was held constant. The subject lay on his or her back in their crib or on the parent’s lap, care being taken that the parent’s face could not be seen by the infant. The face stimulus was of author JC. Her hair was tied back, she wore no make-up or jewelry, and the face was positioned 20 cms above the subject. She adopted a positive, pleasant emotional expression, while remaining silent. Movement of her head was natural, while continuously facing the infant.

The mobile was carefully matched with the face stimulus for 5 factors: (a) Color (‘skin color’). (b) Size and (c) Shape (a ball was used). (d) Contrast (using facial features pasted onto the ball in a scrambled but symmetrical arrangement, following previous studies (Johnson & Morton, 1991)). (e) Dimensionality (to control for a nose-like structure, a 3cm string was attached to the center of the ball, at the end of which was a smaller ball, also matched for ‘skin color’). The mobile itself was attached to a stick 1m in length, and was held above the infant’s head, at the same viewing distance (20 cm). The mobile moved with mechanical motion, since any movement of the larger ball caused the smaller ball to move contingently.

Once the infant was in a state of alert inactivity, a trial began. To be included, an infant had to be looking at the stimulus for at least 3 s. The stimulus was presented for a maximum of 70 s. During this time, a second experimenter filmed the infant’s eye movements. If the infant cried, the trial was suspended, and then restarted so that the total presentation time of the stimulus still amounted to 70 s. If the infant completed .53 s (i.e. 75% of the target time), and then became distressed, the trial was not restarted. Thus, the stimulus was presented for a maximum of 70 s, and a minimum of 53 s. Looking time was calculated as a proportion of total looking time. Care was taken not to film any information that might indicate the sex of the baby.

The videotapes were coded by two judges who were blind to the infant’s sex, to calculate the number of seconds the infants looked at each stimulus. A second observer (independent of the first pair and also blind to the infants’ sex) was trained to use the same coding technique for 20 randomly selected infants to establish reliability. Agreement, measured as the Pearson correlation between observers’ recorded looking times for both conditions, was 0.85, p 5 0.0001.

For each baby, a difference score was calculated by subtracting the percentage of time spent looking at the mobile from the percentage of time they spent looking at the face. Each baby was classified as having a preference for (a) the face (difference score of 120 or higher), (b) the mobile (difference score of 220 or less), or (c) no preference (difference score of between 220 and 120). A 20% cutoff was arbitrarily selected to define a substantive difference in the baby’s interest in the two stimuli. (Selecting other arbitrary cut-offs of 30% or 40% does not affect the results, reported next.)

Table 1 shows the number of babies that fell into each of the 3 categories. A x2 test demonstrated that there was a significant association between sex and stimulus preference (x2 5 8.3, df. 5 2, p 5 0.016). An analysis of adjusted residuals demonstrated that the significant result is due to more of the male babies, and fewer of the female babies, having a preference for the mobile than would be predicted. In other words, male babies tend to prefer the mobile, whereas female babies either have no preference or prefer the real face. This result is supported by considering the mean percentage looking times for male and female babies (see Table 2). A repeated measures ANOVA, comparing percentage looking times for males and females for the face and mobile, found that neither the main effect of sex [F(1, 100) 5 1.03, p . 0.3] or of stimulus type [F(1, 100) 5 0.10, p . 0.7] were significant. There was, however, a significant sex x stimulus type interaction [F(1, 100) 5 5.28, p 5 0.02]. The interaction was investigated using t tests which demonstrated that males looked significantly longer at the mobile than females did (t 5 2.3, df. 5 100, p 5 0.02) and also that females looked longer at the real face than at the mobile (t 5 2.4, df. 5 100, p 5 0.02). The results from the ANOVA were replicated when the age and weight of the baby, duration of trial, and the length of gestation were entered as covariates.

In summary, we have demonstrated that at 1 day old, human neonates demonstrate sexual dimorphism in both social and mechanical perception. Male infants show a stronger interest in mechanical objects, while female infants show a stronger interest in the face. The male preference cannot have simply been for a moving stimulus, as both stimuli moved. Rather, their natural motion differed, the face with biological motion, the mobile with physicomechanical motion. Naturally, these results apply to males and females averaged over a group, and not to all individuals. At such an age, these sex differences cannot readily be attributed to postnatal experience, and are instead consistent with a biological cause, most likely neurogenetic and/or neuroendocrine in nature.

The criticism, in the Journal of Interdisciplinary Feminist Thought:

https://digitalcommons.salve.edu/cgi/viewcontent.cgi?referer=&httpsredir=1&article=1016&context=jift

The theoretical rationale stems from Simon Baron-Cohen’s work on autism (Baron-Cohen, Knickmeyer, & Belmonte, 2005), a condition characterized by social impairment and heightened interest in the physical world. Baron-Cohen views autism, more common in males than in females, as a manifestation of an “extreme male brain”. He further posits that, in general, male cognition reflects an emphasis on analysis and mechanical understanding – what Baron-Cohen terms ‘systemizing.’ Systemizing abilities, in turn, would provide the basis for scientific reasoning. In contrast, female cognition reflects an emphasis on social understanding -- what Baron-Cohen terms ‘empathizing.’ According to Baron-Cohen (2003; Baron-Cohen et al., 2005), these sex differences in ‘systemizing’ and ‘empathizing’ capacities stem from prenatal ‘hardwiring’ in the brain and are present at birth.

Strong support for this theory would be provided by demonstrating that sex differences in these capacities do indeed begin at birth. Unfortunately, although newborns provide an opportunity to examine biological predispositions prior to experience, they are in fact very difficult to study. They drift among various states of consciousness in an unpredictable manner, going from an alert state to crying to sleeping within a very short time. Their attention spans are extremely variable (Fogel, 2001). To address some of these difficulties, standard methodologies are typically used. However, despite an existing body of research on newborn perceptual preferences, Connellan and BaronCohen’s newborn study does not use this methodology, or address the findings from this body of research. We next address these concerns in an in-depth critique of Connellan et al.’s (2000) study.

Connellan et al. assume that sex differences in infants’ interest in ‘social’ versus ‘mechanical’ stimuli are precursors to future sex differences in empathizing and systemizing abilities. They therefore compared one-day-old girls’ and boys’ interest in ‘social’ versus ‘mechanical’ stimuli. The social stimulus was the real, live, face of the first author of the study, Jennifer Connellan. The mechanical stimulus was a mobile – a face-sized ball composed of various facial features that were haphazardly arranged. The stimuli were presented sequentially to infants who were tested at their mothers’ bedsides or in the neonatal nursery, lying on their backs in their cribs or held in their parents’ laps. Differences in time spent looking at the face or the mobile were considered indices of preferences.

This study is fraught with methodological problems:

Validity: The investigators assumed that one-day-old newborns’ preferences for faces or mobiles reflect social versus mechanical intelligence. However, it is not clear that looking at a face or a mobile does, in fact, reflect later social or mechanical abilities. Furthermore, even in newborns, it is unclear that a face represents a ‘social’ stimulus. Recent studies suggest that newborn preferences for face patterns compared to other patterns represent a perceptual bias for top-heavy patterns, rather than preferences for faces per se (Macchi-Cassia, Turati, & Simion, 2004). Indeed, other studies indicate that newborns actually prefer looking at these kinds of patterns to looking at real faces (Simion, Macchi-Cassia, Turati, & Valenza, 2001). It is not until 3 months of age that infants prefer real faces to face-like (top heavy) patterns (Turati, Valenza, Leo, & Simion, 2005).

Confounds/lack of control: The face and mobile differed on several crucial variables. The face in fact consists of several properties. It consists of movement and expressions, and is attached to a live person who exudes warmth and odors. As all of these vary together, it is impossible to know which of these dimensions underlie any preferences for the ‘face.’ Again, it is not clear the preferences for the face can be attributed to its ‘social’ dimension. Furthermore, infants were tested in different settings. Slight changes in the settings could lead to differences in the perception of the two stimuli, making them appear more or less top heavy under different conditions. Preferences for one stimulus or the other could simply reflect these different perspectives.

Experimenter expectancies: A striking design flaw is that the face stimulus was that of the researcher herself. Experimenter and subject expectancies are well-documented, and require stringent controls. In this case, the researcher could unconsciously move her face or tilt her head in ways that increase its salience, which is particular problematic given that in many cases she was aware of the newborn’s sex (Edge, 2005, Simon Baron-Cohen).

Operational definition of the dependent variable: Curiously, the dependent variable itself was incorrectly defined. The authors report that “looking time was calculated as a proportion of total looking time” (p. 115). However, examination of the results indicates that looking time to each stimulus was actually compared to the presentation time for each. In order to interpret findings, the definition of the dependent variable must be clear and precise.

Statistics. Finally, we noticed a serious mistake in the number of degrees of freedom used in the statistical analysis that compared girls’ looking time to the face and mobile. It is very possible that the correct number would not have yielded significant results. Precision in statistical analyses is a fundamental component of scientific investigation. We wonder why this obvious error was overlooked in the peer review process for journal publication. These are all serious methodological problems. There are numerous investigations that adopt a standard research paradigm for investigating perceptual preferences: all infants are tested in the same, controlled setting, stimuli are presented simultaneously to control for the fragility in newborn attention span and fluidity of newborn states, and specific procedures are used to prevent parents and researchers from influencing infants’ behavior. In studies of sex differences, experimenters are blind to the sex of the infants. Finally, the dependent variable is measured more precisely – in terms of actual looking time (the difference in the amount of time that infants looked at one stimulus compared to another) rather than relative looking time. It is unclear why the only study of sex differences in perceptual preferences in newborns ignored these standard procedures.

Findings and Conclusion

The only significant findings were that boys looked at the mobile more than girls did, and girls looked at the face more than they looked at the mobile. Much was made of these findings in the paper’s conclusion: the authors state that they “demonstrate beyond a reasonable doubt [italics added] that these [sex] differences are in part biological in origin.” (p. 114). This is a strong claim, with serious implications. The methodological problems already provide reasonable doubt. We next show that the findings themselves are equivocal in regard to this claim. The authors conclude that “We have demonstrated that at 1 day old, human neonates demonstrate sexual dimorphism in both social and mechanical perception. Male infants show a stronger interest in mechanical objects, while female infants show a stronger interest in the face” (Connellan et al., 2000, p. 116). However, this conclusion incorporates comparisons that were not significant: Boys did not look at the mobile more than the face; they only looked at the mobile more than girls did, and girls did not look at the face more than boys did; they only looked at the face more than the mobile. The title of the paper, “Sex Differences in Human Neonatal Social Perception,” is in fact misleading, as the findings indicated no sex difference in looking time to the face.