Method
The two gaps
The gender composition of high-achieving STEM applicants reflects two margins: the gender composition of the high-achieving group and the rates at which women and men in that group apply to STEM.
1. The quantity being split
Fix a setting and a threshold — say, the top 10% of the academic performance distribution. Two facts describe the STEM applicant pool at that threshold. First, who is in the high-achieving group: the share of students above the bar who are women. Second, what they do: among those students, how likely a woman and a man each are to rank a STEM program first.
The first is the pipeline gap, defined as the female-minus-male difference in representation among students above the threshold. The second is the choice gap, defined as the female-minus-male difference in the share of high-achieving students who rank STEM first or anywhere on the list. A percentage point (pp) is an absolute difference in probabilities: a gap of −20pp means that if 30% of qualifying men rank STEM first, 10% of qualifying women do.
2. The identity
Let Tg denote the number of students of gender g in the top group and sg the share who rank the field first. The number of high-achieving applicants is Ag = Tgsg. Taking logs of the ratio between men and women gives
ln(AM / AF) = ln(TM / TF) + ln(sM / sF) = P + C
The first term is the pipeline component P; the second is the choice component C. This is an accounting identity with no residual. Its orientation differs from the percentage-point figures: those report female minus male, whereas positive values of P or C favor men.
The two components are directly comparable in this form. The underlying percentage-point gaps are not: one is a difference in composition shares and the other a difference in application rates.
3. From a component back to a share
The male share among high-achieving applicants is Λ(P + C), where Λ(z) = 1 / (1 + e−z). A value of zero maps to a male share of 50%.
Setting P to zero gives Λ(C): the male share that would obtain if the top group were gender balanced, holding the observed choice component fixed. Setting C to zero gives Λ(P): the share that would obtain if women and men applied at equal rates, holding the observed pipeline component fixed.
4. Why deferred acceptance matters
The obvious objection to reading application lists as preferences is strategy. If a student expects to be rejected from her first choice, she might not list it. Then a list reveals a mixture of what she wants and what she thinks she can get, and a gender difference in the lists could be a gender difference in confidence about admission rather than in what students want to study.
All ten systems assign places through variants of deferred acceptance: students submit a ranked list, programs rank students by a performance score, and the algorithm repeatedly tentatively assigns and bumps applicants until the assignment is stable. Under this class of mechanism, ranking programs in true preference order generally gives applicants little incentive to misreport. This argument is weaker where lists are short, applications precede scores, or the mechanism is iterative; the paper treats these features as qualifications rather than assuming the systems are frictionless.
- All ten systems rank students on academic performance and assign places through variants of deferred acceptance.
- Deferred acceptance gives applicants little incentive to misreport, so the submitted list is informative about the ordering of programs they were willing to enter.
- List-length caps range from two programs in Brazil to a hundred in Taiwan, and Greece imposes no numerical limit at all.
5. Why high achievers
The analysis concentrates on students near the top of each performance distribution because that is where the constraints that usually confound this comparison bind least. Students in the top decile are the applicants most likely to gain admission to and succeed in selective STEM programs. Restricting attention to this group reduces, but does not eliminate, differences in admission prospects.
The top-10% cutoff is a reporting convention. The appendix repeats the analysis at the top 5%, 20%, and 30% and across the full performance distribution; top-1% estimates are also available for nine settings.
6. What the design does not establish
The institutional design reduces the scope for several mechanical explanations, but it does not identify the process that generates the observed rankings.
Two further limits are worth stating plainly. The decomposition is descriptive — it allocates an observed gap between two margins, and does not estimate the effect of any intervention. And the comparability of the ten settings rests on the construction choices documented on each country page, including the performance measure, the sample definition, and, in Chile, a denominator conditioned on students who submit a list.
The choice component compares women and men within the top group rather than at an identical score. Differences in their positions within that group may therefore contribute to the component. Its persistence at narrower thresholds shows that such composition cannot account for the whole component; it does not show that composition makes no contribution.