Evidence explainer

Women's, men's, and reproductive health

The Women's Health Research Gap Is More Than Enrollment

Counting women into trials was necessary. It was not sufficient, and the gap that remains is largely about which questions get asked at all.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. The historical correction changed who could enter studies
  2. Denominator choice can make representation look better or worse
  3. Enrollment without analysis leaves a partial gap
  4. Preclinical design shapes the questions that reach trials
  5. Sex and gender answer different parts of the problem
  6. Pregnancy creates evidence and ethics tension
  7. Menopause and aging reveal a lifespan gap
  8. Conditions can be unique, more common, or different
  9. Relevant outcomes may differ from traditional endpoints
  10. Closing the gap needs an audit trail
  11. References

The women's health research gap is often described as too few women in clinical trials. That history matters, and enrollment has improved. NIH reports that women now account for roughly half of participants in NIH-supported clinical research overall. The remaining gap is not captured by one percentage.

A portfolio can reach numerical parity while leaving major questions unanswered. Women may be concentrated in some conditions and scarce in others. Trials may be too small to estimate different effects. Pregnant and breastfeeding women are commonly omitted. Female cells and animals may be absent upstream. Conditions unique to women, more common in women, or experienced differently can receive inadequate scientific attention.

The historical correction changed who could enter studies#

For much of modern drug development, women of childbearing potential were often excluded from early trials, and research was frequently conducted in male animals or without reporting sex, and concern about fetal harm after historical drug disasters contributed to protective policies that became broad exclusion.

Protection from research risk is ethically essential. Exclusion also has a cost. A medicine approved on evidence gathered mainly in men can end up in your hand years later, without adequate knowledge of dose, adverse effects, interactions, menstrual or hormonal influences, pregnancy, or lactation.

The NIH Revitalization Act of 1993 established inclusion requirements for women and minority groups in NIH-supported clinical research. FDA policy evolved to encourage participation and analysis by sex. These changes improved representation, but they did not automatically rebalance every field or create enough events for reliable subgroup estimates.

Denominator choice can make representation look better or worse#

The relevant benchmark is not always 50 percent. A trial of a condition affecting women and men equally may seek balance, while a trial of a condition more common in women should reflect the population likely to use the intervention. A prostate or ovarian study has a different target by anatomy and indication.

Report representation relative to disease burden and expected use, not only the share of all participants. Trial phase matters too. Women may enter later efficacy studies while remaining less represented in dose-finding or device studies. Country and site selection can change age and comorbidity. Intersection matters. An overall female share can conceal low participation among older women, women from underrepresented racial or ethnic groups, women with disabilities, rural residents, people with limited English proficiency, or those with caregiving constraints.

Enrollment without analysis leaves a partial gap#

A trial can include women but remain unable to estimate whether benefit or harm differs. Interaction analyses require adequate sample size and event counts. Simply testing statistical significance within women and within men is wrong: one subgroup can be significant and the other not even when effects are compatible.

The correct question is whether the treatment contrast differs between groups, estimated with an interaction and uncertainty. Most trials are not powered for modest subgroup effects. A null interaction may therefore be inconclusive rather than proof of identical effects.

Prespecification reduces selective reporting. Researchers should state the rationale, scale, and direction of a proposed difference. Pooling individual-participant data across trials can improve precision. Pharmacokinetic studies may detect dose-related differences that an outcome trial cannot. Sex-disaggregated reporting remains useful even when formal interaction estimates are imprecise, because it shows you who enrolled, who withdrew, and where the evidence is thin.

Preclinical design shapes the questions that reach trials#

NIH's sex-as-a-biological-variable policy asks funded researchers to consider sex in vertebrate animal and human studies, from question and design through analysis and reporting. The goal is not to force every study into an underpowered comparison. It is to make inclusion, exclusion, and pooling deliberate.

Cell sex can matter through chromosomes, gene dosage, hormones, epigenetics, and tissue history. Animal physiology can vary with sex, strain, age, housing, and reproductive stage. Estrous cycling is sometimes cited as a reason to use only males, yet male animals also vary biologically.

Studying both sexes does not mean every observed difference is meaningful. Replication, mechanism, multiple testing, and effect size matter. A result in mice should not be assigned directly to women or men without human evidence.

Sex and gender answer different parts of the problem#

Sex-related variables can include chromosomes, reproductive anatomy, hormones, pregnancy, and menopause. Gender-related variables can include social roles, identity, and discrimination. They can include occupation, caregiving, behavior, and health-care interactions. Their definitions and legal collection rules vary by study and jurisdiction.

They interact. Occupational patterns can change chemical or injury risk. Caregiving can affect sleep and appointment access. Hormonal transitions can influence symptoms and pharmacology. Clinician interpretation can change diagnosis.

Using a binary sex field as a substitute for all these pathways produces vague science. Measure the factor relevant to the hypothesis when possible. A study of drug clearance may need kidney function, body composition, enzymes, hormones, and pregnancy status; a study of delayed diagnosis may need symptoms, referral pathways, insurance, caregiving, and clinician behavior.

Pregnancy creates evidence and ethics tension#

Pregnancy changes plasma volume, kidney filtration, body composition, drug-metabolizing enzymes, and placental transfer. The same dose can produce different concentrations across gestation. Illness and undertreatment can harm both pregnant person and fetus, while an inadequately studied medicine can create unknown risk.

FDA notes that pregnant women have historically been excluded from drug-development trials and that inclusion can be scientifically and ethically appropriate in some circumstances. The 2025 draft ICH E21 guidance proposes a structured approach to including and retaining pregnant and breastfeeding women and generating evidence for decisions.

Inclusion should be based on the condition, existing reproductive toxicology, and prior human data. It should be based on gestational timing, alternatives, and monitoring. It should be based on consent and prospect of benefit. Pregnancy registries and postmarket studies help but often face delayed enrollment, missing comparators, and incomplete outcomes.

Breastfeeding research needs information about transfer into milk, infant dose, oral absorption, infant outcomes, and effects on milk production, though a detectable concentration is not the same as clinical harm, and absence of detection depends on assay sensitivity and timing.

Menopause and aging reveal a lifespan gap#

Women's health research is sometimes reduced to reproduction. Menopause affects symptoms, bone, and cardiometabolic risk. It affects sleep, sexual health, and treatment choices across decades. Older women also have multimorbidity, polypharmacy, frailty, and caregiving conditions that younger trials may not represent.

Timing matters. A result in early postmenopause may not transport to treatment initiated much later. Outcomes and baseline risk change with age. Trials need enough duration to assess benefits and harms relevant to chronic use.

The National Academies' 2024 report called for a broader vision of women's health across the life course and identified research gaps across selected conditions and systems. Portfolio analysis asks not only who enters studies, but which conditions receive questions, methods, and sustained funding.

Conditions can be unique, more common, or different#

Some conditions are specific to reproductive organs or pregnancy. Others, including several autoimmune diseases, migraine, osteoporosis, and some pain conditions, occur more often in women. Cardiovascular disease can present and be recognized differently, while average lifetime risk and subtype patterns vary.

When you read that a condition “affects women differently”, ask which dimension is meant: incidence, symptoms, or test performance. It could be treatment response, adverse effects, access, or outcome. Broad claims that women are simply atypical preserve a male default. Studies should define the target population and measure the pathway.

The health-screening guide provides a clinical overview, while why diagnosis wording matters examines classification and communication.

Relevant outcomes may differ from traditional endpoints#

Research agendas should include function, pain, fatigue, fertility goals, sexual health, caregiving burden, return to work, mental health, and quality of life where relevant, and these outcomes are not “soft” merely because they are reported by patients.

Measurement tools require validation across languages, ages, and populations. A scale developed in one narrow group may miss symptoms or meanings elsewhere. Core outcome sets developed with patients can improve comparison without erasing individual priorities. Harms need adequate ascertainment. Menstrual changes, pregnancy outcomes, sexual effects, and symptoms that are stigmatized or normalized can be underreported unless asked respectfully and specifically.

Closing the gap needs an audit trail#

Funders can map burden, uncertainty, and portfolio allocation. Investigators can justify population and model choices, plan analyses, and publish null subgroup findings. Regulators can request representative evidence and postmarket follow-up. Journals can require disaggregated reporting and reject biological stories unsupported by design.

Health systems can test whether evidence transports to the people receiving care. Patients and community partners can shape priorities before a protocol is fixed. No single mandate closes every gap, but each actor can make missing evidence visible.

Progress should be measured in decisions improved, not only participants counted. The objective is evidence that helps you weigh benefit and harm in the real conditions of your own life. That is a stricter standard than parity on a recruitment chart, and a more useful one.

References#

  1. National Academies, A New Vision for Women's Health Research
  2. NIH sex as a biological variable
  3. NIH inclusion policy
  4. FDA women in clinical trials
  5. FDA clinical trials in pregnant women
  6. FDA draft E21 guidance

For your own health, talk with your clinician.*

Questions and answers

Are women still absent from most clinical research?

Women now make up about half of participants in NIH-supported clinical research overall, but representation and analyzable evidence vary greatly by condition, trial phase, age, and pregnancy status.

Why is equal enrollment not enough?

A trial can enroll many women yet lack power for sex-specific effects, use irrelevant outcomes, omit pregnancy or older age, or fail to report results by sex.

What does sex as a biological variable mean?

It means considering whether biological sex may affect the research question, design, analysis, and reporting in cell, animal, and human studies, with justification when one sex is studied.

Why have pregnant and breastfeeding women often been excluded?

Concern about fetal or infant harm, liability, and complex physiology has encouraged exclusion, but the result is uncertain dosing and safety when treatment is clinically necessary.

How can research close the gap without assuming every difference is biological?

Measure relevant biological, social, behavioral, and structural factors; prespecify analyses; avoid using sex as a proxy; and report uncertainty rather than assigning a mechanism from a subgroup result.