Dark Mode Light Mode

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.
Join us on a journey where chemistry meets creativity, and the wonders of science unfold. Quench your intellectual thirst with thought-provoking articles that transcend the boundaries of conventional knowledge.

Maternal influences on infant gut microbiome and health

Maternal influences on infant gut microbiome and health Maternal influences on infant gut microbiome and health


Study design

The samples for this study were obtained from the Lifelines NEXT (LLNEXT) cohort, a birth cohort designed to study the effects of intrinsic and extrinsic predictors of health and disease in a four-generation design23. LLNEXT is embedded within the Lifelines (LL) cohort study, a prospective three-generation population-based cohort study recording the health and health-related aspects of 167,729 individuals living in the northern Netherlands56,57. From 2016 to 2023, 1,452 pregnant women were recruited for LLNEXT, along with their partners and infants, and monitored up to at least 1 year after birth. Several biomaterials, including faeces, breast milk and vaginal swabs, were collected from participants. Data on medical, social, lifestyle and environmental factors were collected through questionnaires at 14 different timepoints and through connected devices. The LLNEXT study was approved by the Ethics Committee of the University Medical Center Groningen (UMCG), document number METC UMCG METc2015/600. Written informed consent forms were signed by the participants or their parents/legal guardians. The present study included the first 714 mother–infant pairs recruited between 2016 and 2019.

Predictors

Extensive predictor data were collected for all participants through questionnaires and medical records from the LLNEXT and LL cohort studies (if the parents were part of LL). Information from several sources were combined to ensure completeness. Data regarding predictors included information from mothers from before pregnancy, during pregnancy, at delivery and after pregnancy. Information from infants were obtained during birth and over the first year after birth. Predictors were investigated both cross-sectionally (for example, static predictors such as delivery mode; n = 316 predictors; Supplementary Table 1) and longitudinally (for example, dynamic predictors such as feeding mode; n = 158 predictors; Supplementary Table 2).

Pre-pregnancy predictors

Maternal pre-pregnancy BMI (kg m−2) data were obtained from the LL cohort study and complemented with medical records of LLNEXT to obtain BMI measurements for mothers up to 8 years before conception.

Maternal pre-pregnancy smoking history was classified as ‘yes’ if the mother reported smoking consistently for at least one full year at any point in her life. This period of smoking in the woman in LLNEXT occurred between the ages of 10 and 29, with a median age of 15 years.

Pregnancy and delivery predictors

Maternal education and income were reported at week 18 of pregnancy (P18). Mother education was categorized as follows: elementary to lower-secondary education (including incomplete and completed special primary education, primary or pre-vocational education and general secondary education), upper-secondary education (including vocational or secondary vocational education and higher general and pre-university education); or tertiary education (including higher professional education and university)58. Maternal income per month was derived using data from the Dutch Social and Cultural Plan59 and categorized as low (<€800–1,800), middle (€1,800–2,400) and high (>€2,400) income. Questionnaires related to maternal reproductive and dental health, as well as substance use were also completed at P18.

Maternal food preference and avoidance questionnaires, comprising 94 items, were completed at P12 and P32 and were developed by Wageningen University & Research (WUR). Food preference and avoidance items were evaluated using a Likert scale ranging from 0 (strong aversion) to 10 (strong preference). Given that these items were highly correlated (r > 0.8) at both timepoints, the questionnaires were combined, and missing data were imputed using a linear regression approach of one timepoint to another.

Delivery predictors were recorded through comprehensive medical records (birth cards completed by midwives or obstetricians at home or in the hospital) and questionnaires, which included details on delivery mode, place of delivery, duration of delivery, medication administered during pregnancy and pregnancy complications.

Postpartum predictors

At M2, maternal dietary intake was assessed using an online self-administered, semi-quantitative Dutch food frequency questionnaire (FFQ) comprising 177 food items (Supplementary Table 72). The FFQ was quality-checked by trained research dieticians at WUR. FFQs referred to the previous month, and mothers reported consumption frequencies from never to 7 days per week. Portion sizes were estimated in g per day using natural portions and household measures according to the ‘Measures, Weights and Codes’ booklet60. Daily energy intake (kcal per day) was calculated by multiplying consumption frequency by portion size and energy content from the Dutch food composition table (NEVO) of 2006 (Dutch Food Composition Database (NEVO), RIVM61). Mothers reporting total caloric intakes <500 kcal per day or >4,000 kcal per day were excluded (n = 6). Food items were grouped into 28 broader food categories (Supplementary Table 72), each named according to the most abundant food item(s) within that category. Moreover, an alternative Mediterranean diet quality score was derived62. To derive dietary patterns, we performed unsupervised hierarchical clustering on the food groups, measured in g per day using Euclidean distances, followed by complete-linkage clustering. Various clustering heights were tested and a dendrogram was used to identify the optimal cut height for the branches. A cut height of 26 was selected, and centroids were extracted to represent the mean consumption of all variables within a cluster (n = 10) for each mother. This was performed using the dendextend package (v.1.17.1)63 in R.

Infant predictors

Infant stool consistency was measured using the Brussels Infants and Toddlers Stool Scale (BITSS64) at 2 weeks and 1, 2, 3, 6, 9 and 12 months old. Infant stool characteristics and gastrointestinal (GI)-related symptoms were derived from the validated ROME IV questionnaire65 at 1, 2, 3, 6, 9 and 12 months old. The 13-item Infant Gastrointestinal Symptom Questionnaire (IGSQ) was used to assess parent-perceived GI functioning and distress, along with feeding tolerance in infants at 3 and 12 months old. Items were scored according to the user guidelines66, and a composite IGSQ score was derived as instructed. Scores were dichotomized at each timepoint according to the population median (as scores were right-skewed): infants with scores at or below the population median were considered to experience no-to-little GI distress and those with scores above the population median were considered to experience medium-high GI distress.

Infant feeding mode and dietary intake were obtained from medical records (that is, birth cards) and FFQs. FFQs were completed by parents or guardians and were specifically designed for Dutch infants, covering items over the previous month at 6 (49 items), 9 (66 items) and 12 (66 items) months old (Supplementary Table 37). These FFQs underwent quality checks by trained research dieticians at WUR. Feeding mode at birth (initial feeding mode infants had directly after birth), which included breastfeeding, mixed feeding (defined in this study as any ratio of breastfeeding and formula feeding (FF)) and FF, was obtained from birth cards. Feeding mode after birth was investigated longitudinally, and missing data were imputed, where possible, based on feeding mode combinations at 2 weeks to 3 months old. For example, infants with the following feeding mode combinations (2 weeks–1 month–2 months–3 months): FF–NA–FF–FF, FF–FF–NA–FF, were imputed to be formula-fed infants for all available timepoints from 2 weeks to 3 months old. Complementary feeding was derived from food items of the 6, 9 and 12 month FFQs according to the CDC definition: “foods or drinks other than breast milk or infant formula (e.g. infant cereals, fruits, vegetables, water)”. Like the mother FFQ at month 2, intake was recorded based on consumption frequencies ranging from ‘never’ to ‘7 days per week’. Portion sizes were estimated in g per day using the ‘Measures, Weights and Codes’ booklet67, and daily energy intake (kcal per day) was calculated using information from the Dutch food composition table (NEVO) of 200661. Owing to challenges in accurately quantifying breast milk intake (in grams and energy), total daily grams and energy intake were calculated only for complementary foods, excluding breast and formula milk. Macronutrients were converted to energy percentages, where carbohydrates and protein contain 4 calories per gram, fat 9 calories per gram and fibre 2 calories per gram. Food items were grouped into 19 groups at the 6 months old and into 22 groups at 9 and 12 months old (Supplementary Table 37). Considering the dietary patterns prevalent in the Netherlands, food items related to dairy and fruit were investigated as both individual and grouped items. These items were assessed based on their consumption (yes/no) and weighted frequencies (that is, with ‘1 day per week’ being equivalent to 0.04, given that 1 month was considered to comprise 28 days). Infant dietary patterns at 6, 9 and 12 months old were derived by factor analysis (FA), using the extraction method of principal components on all standardized FFQ items (in g per day). Suitability for conducting FA was confirmed by Bartlett’s test of sphericity and the Kaiser–Mayer–Olkin test, where Kaiser–Mayer–Olkin values for all items were >0.5. To determine the number of dietary patterns to retain, eigenvalues ≥ 1.5 were considered relevant and the ‘elbow’ of the scree plot was investigated. This was followed by orthogonally rotating (using the varimax rotation method) the components. For each component, food items with loadings |>0.35| were considered to contribute significantly to the pattern. Component scores were calculated with the regression method68. Bartlett’s test of sphericity and the Kaiser–Mayer–Olkin test were conducted using the psych package (v.2.2.9)69 in R, and FA was implemented using the factanal function of the stats package (v.4.2.1)70 in R.

Infant anthropometric measurements, which included weight (kg), length (cm) and head circumference (cm), were obtained from medical records, growth charts and questionnaires from birth to 12 months old. Missing data were estimated, from birth to 12 months old, according to third-degree polynomial regressions of the separate weight, length and head circumference measurements. s.d. scores for all measurements were generated using the Growth Analyzer Research Calculation Tool (v4.1), using data from the Fifth Dutch Growth Study of 2008–2009 as a reference71.

Infant development was assessed based on the 30-item ASQ at 6 months old (ASQ-6) and 12 months old (ASQ-12). The total ASQ score (range: 0–300) and its five composite domains (range: 0–60 each) of communication, gross motor, fine motor, problem solving and personal–social skills were derived according to the ASQ-3 User’s Guide72. Missing data were handled according to the ASQ-3 user guide. Infant development was classified as either development_typical or development_delay based on relaxed and strict thresholds for each domain, as specified in the guide. ASQ-6 scores were derived for infants aged 5–7 months. ASQ-12 scores were derived for infants aged 11–13 months.

Infant health was derived and complemented from general infant health questionnaires (including skin irritation, eczema and diseases) at 4 months and 12 months old; food allergy questionnaires at 4, 6, 9 and 12 months old; and the International Study of Asthma and Allergies in Childhood (ISAAC) questionnaire73 at 12 months old. Infant atopic dermatitis (in this paper, referred to as eczema) was assessed by research nurses at 3 and 12 months old using the SCORring Atopic Dermatitis (SCORAD) approach74. SCORAD evaluates the area of affected skin, itching, sleeplessness and other symptoms, generating a score with a maximum of 83 that can be used to aid in diagnosing infants with eczema. In our study, infant eczema was defined in three ways: ‘infant_health_eczema_questionnaire’ (based on parent-completed questionnaires), ‘infant_health_eczema_diagnosis_strict’ and ‘infant_health_eczema_diagnosis_relaxed’. For the latter two definitions, infants were considered to have eczema if they used eczema medication between 2 weeks and 12 months old, parents reported that their infant had eczema, or the objective SCORAD score was >0. For the strict definition, infants without eczema (controls) were defined as infants who did not use eczema medication, whose parents reported that their infant did not have eczema and whose objective SCORAD scores were 0. For the relaxed definition, controls identified by the strict definition were complemented with predictors derived from the first definition (infant_health_eczema_questionnaire).

Infant crying behaviour was assessed at 2 weeks and 1, 2, 3, 6, 9 and 12 months old by an 11-item crying questionnaire that was previously validated in Dutch infants75.

Infant sleeping behaviour was assessed at 3, 6 and 12 months old with the 13-item brief infant sleep questionnaire (BISQ)76.

Infant medication use was reported by parents at 2 weeks and 1, 2, 3, 6, 9 and 12 months old and complemented with medication use reported in other LLNEXT questionnaires. The in-house-developed tool SORTA (System for Ontology-based Re-coding and Technical Annotation)77 was used to semi-automatically match textual medication predictors with Anatomical Therapeutic Chemical classification codes, followed by a manual curation of any unmatched medications. Infant medications at all timepoints were longitudinally investigated and updated to refer to the infant having ever received the specified medication up until the timepoint assessed.

Other predictors across pregnancy, delivery and/or postpartum

Data on family pets were combined from several maternal (P18, P32 and 1 month and 4 months postpartum) and infant (4 and 12 months old) questionnaires. Family living situation predictors were combined from maternal questionnaires at P32 and 4 months postpartum and included ‘house’, ‘farm’, ‘flat’ or ‘other’ types of living situations.

Maternal stool consistency was measured using the Bristol Stool Scale at P12, P28, birth and 3 months postpartum. The validated ROME III questionnaire78 was completed at P28 and 3 months postpartum and was used to characterize functional gastrointestinal disorders (FGIDs). Mothers were classified as having either no FGIDs or irritable bowel syndrome, functional constipation or functional bloating. Stool frequency characteristics were also derived from the ROME III questionnaire. Moreover, stool diaries were completed by mothers at P28 and 3 months postpartum to assess maternal gastrointestinal health.

Maternal medication use was self-reported at P28, birth and 3 months postpartum. Similar to infants, the in-house SORTA tool was used to semi-automatically match textual medication-related predictors with Anatomical Therapeutic Chemical codes, followed by the manual curation of unmatched medications.

The 12-item Long-term Difficulties Inventory (LDI)79, which assesses maternal stress in the past year, was completed by mothers at 1 and 10 months postpartum. Mothers were asked to rate stressful situations related to work, home, family and other as not stressful (score: 0), somewhat stressful (score: 1) or very stressful (score: 2). Items were summed and averaged to generate a pregnancy and post-pregnancy LDI stress score.

Correlations between predictors

Spearman correlations were computed for all predictors (Supplementary Tables 73–76). To avoid collinearity in subsequent analyses, we excluded predictors with correlation coefficients >0.7 or <−0.7 with FDR < 0.05, including certain mother and infant diet variables, infant growth measurements and ASQ predictors.

Sample collection

Faecal samples

Overall, 1,587 maternal and 2,939 infant faecal samples were collected and successfully sequenced from 714 mother–infant pairs across 10 timepoints. Mothers collected their samples during pregnancy at P12 (n = 414) and P28 (n = 406), during delivery (n = 268) and 3 months postpartum (n = 499; Supplementary Tables 12 and 13). Parents/guardians collected infant faecal samples (n = 2,939) at 2 weeks old (n = 331), and 1 (n = 466), 2 (n = 497), 3 (n = 553), 6 (n = 346), 9 (n = 337) and 12 (n = 409) months old, with an average of four samples collected per infant. The exact timeframe within which the infant samples collected in each of these categories are as follows: 2 weeks old, day 3 to day 22; 1 month old, day 23 to day 45; 2 months old, day 46 to day 77; 3 months old, day 78 to day 143 (4.69 months); 6 months old, day 148 (4.86 months) to day 232 (7.6 months); 9 months old, day 244 (8.01 months) to day 329 (10.81 months); and 12 months old, day 335 (11.01 months) to day 436 (14.32 months). Meconium was also collected from infants, but the very low microbial biomass in these samples meant our attempts to isolate viable microbial reads were unsuccessful, as previously reported20. These samples were therefore not included in the current study. As previously described, parents used stool collection kits provided by the UMCG and froze the faecal samples at home at −20 °C within 10 min of stool production20. Frozen samples were collected by UMCG personnel, transported to the UMCG in portable freezers and stored in −20 °C (short-term) or −80 °C (long-term) freezer until DNA extraction.

Human breast milk samples

Mothers were asked to pump breast milk with their usual breast pump equipment and without specific cleaning of the breast tissue. They were instructed to pump milk from one breast, preferably the right, from the second feeding after midnight with a time interval of at least 2 h since the last feed from that breast. Mothers homogenized milk samples by gentle shaking and, using plastic dropper pipettes (H10041, MLS), prepared 2 ml aliquots in cryotubes (122279, Greiner Bio-One). Milk samples were stored at −20 °C in home freezers, transported to the laboratory in transportable freezers, then stored at −20 °C (short-term) or −80 °C (long-term) until analysis.

Vaginal samples

Women were asked to collect a vaginal swab close to birth. This was placed by mothers in a Power Bead Solution (Qiagen) and stored at −20 °C (short-term) or −80 °C (long-term) until analysis.

Microbiome processing and profiling

Faecal DNA extraction

Microbial DNA was isolated from 0.2–0.5 g faecal material using the QIAamp Fast DNA Stool Mini Kit (Qiagen) and the QIAcube (Qiagen), according to the manufacturer’s instructions, at the Institute for Clinical Molecular Biology, Kiel, Germany, with a final elution volume of 100 μl. DNA eluates were stored at –20 °C.

Human breast milk and vaginal DNA extraction

DNA was isolated from 3.5 ml breast milk and from vaginal swabs using the DNeasy PowerSoil Pro Kit (47016, Qiagen), as described previously80,81. In brief, breast milk samples were centrifuged at 13,000g for 15 min at 4 °C. Fat and whey were removed. Cell pellets were resuspended in 800 µl solution CD1, transferred to PowerBead Pro tubes and vortexed briefly. Vaginal swab samples were thawed at 4 °C for 1 h. Each sample was then vortexed for 3 min in 500 µl of PowerBead Solution. The liquid was transferred to PowerBead Pro Tubes, the total volume adjusted to 800 µl using CD1 solution and then vortexed briefly. Prepared milk and vaginal samples were incubated at 65 °C for 10 min and subsequently bead-beat at 5,000 rpm at 4 °C for 45 s on a Precellys Evolution tissue homogenizer (Bertin Instruments). The samples were then centrifuged at 15,000g for 1 min at 4 °C, and a 600 µl sample was used for automatic DNA extraction on Qiacubes (Qiagen) using the ‘DNeasy PowerSoil Pro Kit with Inhibitor Removal Technology Protocol’. DNA was eluted in 50 µl and stored at −20 °C.

Genomic library preparation and sequencing

Faecal, vaginal and breast milk microbial DNA samples were sent to Novogene, Cambridge, UK for genomic library preparation and shotgun metagenomics sequencing. Sequencing libraries were prepared using the NEBNext Ultra DNA Library Prep Kit or the NEBNext Ultra II DNA Library Prep Kit (depending on the sample DNA concentration), and sequenced using HiSeq 2000 or NovaSeq 6000 sequencing with 2 × 150 bp paired-end chemistry (Illumina), as previously described20.

Profiling of the gut, vaginal and breast milk microbiome

Bioinformatic analysis was performed in-house using our bioinformatic pipeline (https://github.com/GRONINGEN-MICROBIOME-CENTRE/gmc-mgs-pipeline). In brief, adapters were first trimmed from the gut microbiome sequencing reads with BBDuk (v.39.01)82 and then quality trimmed using KneadData (v.0.10.0)83. Thereafter, the KneadData-integrated Bowtie2 tool (v.2.4.2)84 was used to remove reads that aligned to the human genome (GRCh38/hg38), and the quality of the processed data was examined using the FastQC toolkit (v.0.11.9)85. Taxonomic composition of metagenomes was profiled using the MetaPhlAn4 tool with the MetaPhlAn database of marker genes mpa_vJan21 and the ChocoPhlAn Species level Genomic Bin (SGB) database (202103)86. Bacterial strain haplotypes were generated using StrainPhlAn486. This method is based on reconstructing consensus sequence variants within species-specific marker genes and using them to estimate strain-level phylogenies. It considers only the dominant strain of species and therefore misses overlap in secondary strains. We profiled the abundance of microbial metabolic pathways in gut microbiomes using HUMAnN (v.3.6)83.

SGB filtration and data transformation for gut microbiome analysis

For the downstream analysis of SGBs, maternal and infant gut microbiome data were first separated. For mothers, we set a relative abundance cut-off of 0.001% and a prevalence cut-off of 30%, resulting in 322 SGBs for further analysis. For infants, we set a relative abundance cut-off of 0.1% and a prevalence cut-off of 10%, resulting in 105 SGBs for further analysis. CLR transformation was then applied at the SGB level and at higher taxonomic levels. All microbial taxa, regardless of taxonomic level, were CLR-transformed using the geometric mean of the relative abundance of microbial species as the CLR denominator. As CLR transformation cannot be applied to zero values, zeros were adjusted by adding half of the lowest non-zero value.

MetaCyc pathway filtration and data transformation

We filtered metabolic pathways based on a prevalence of >30% and a minimum relative abundance of 0.005% for both mothers and infant gut microbiome profiles separately, resulting in 289 pathways for infants and 171 pathways for mothers. Transformation for MetaCyc pathways was performed using ALR transformation with the geometric mean of species abundances as the denominator. As the external denominator was used, we applied additional treatment for jitter introduced to zero abundances by subtracting the geometric mean. For each pathway, all values that were zero before the transformation are made equal and placed below the smallest non-zero value.

GBM filtration and data transformation

GBMs comprise curated modules of microbial pathways involved in the metabolism of molecules with the potential to interact with the human nervous system37. We performed this analysis selectively in infants including 34 GBMs present in over 30% of infants in the analysis. Transformation for GBMs was performed using ALR transformation with the geometric mean of species abundances as denominator.

CAZyme filtration and data transformation

CAZymes were annotated using Cayman87, and a prevalence filter of 30% was applied separately to maternal and infant samples. Hand-annotated substrate specificity per CAZyme was retrieved87. Transformation for CAZymes was performed using ALR transformation with the geometric mean of species abundances as denominator.

HMO profiling

The concentrations of 24 HMOs were measured in maternal breast milk samples and quantified using ultra-high-performance liquid chromatography as described previously88. To determine maternal Le and Se status, the milk concentrations of HMOs LNFP-II and of 2′FL, respectively were used as proxies.

Associations between timepoint and alpha and beta diversity and species-level composition

To investigate microbial diversity within samples, the microbial alpha diversity of gut microbiome samples, as represented by the Shannon diversity index, was calculated on SGB relative abundances using the diversity function of the vegan package (v.2.7-1)89 in R. To test the effect of timepoint on alpha diversity and each sample, we tested this as a fixed effect in a mixed model using the lmerTest package (v.3.1-3)90 in R (Supplementary Tables 16 and 18). Read depth, sample DNA concentration and batch number were included as covariates and sample ID as a random effect. P12 was used as the reference for mothers and 2 weeks old as the reference for infants. To calculate the beta diversity within and between samples, we used Aitchison distances performed on filtered CLR-transformed data. The association between microbial beta diversity and time was performed using PERMANOVA (adonis2 analysis of the vegan package) with constraint to within-participant sample permutations (10,000) to derive P and R2. The effect of time on CLR-transformed SGB relative abundances in mothers and infants was calculated using timepoint as a fixed effect in a mixed model using again the lmerTest package with the same covariates and references (Supplementary Tables 17 and 19).

Early-life bacterial composition clustering analysis

To define early-life clusters, 2-week-old samples were used. A prevalence cut-off of 10% on a minimum relative abundance of 0.1% was applied to focus on bacterial species at SGB level, which were prevalent in the population with medium to large relative abundances (33 out of 512 species). Using the filtered taxonomy table, we used the vegan package in R to calculate Bray–Curtis dissimilarities, which were governed by the few dominant species at 2 weeks old. Hierarchical clustering (complete-linkage) was then performed on the dissimilarity matrices. We identified the optimal number of clusters by applying a cut-off that optimized the Calinski–Harabasz index91. This was done using the as.clustrange function from the WeightedCluster package (v.1.6-4)92 in R, with a maximum number of 20 clusters.

We next assessed the community composition of the clustered samples at late timepoints (6, 9 and 12 months old) and attempted to predict the early 2-week-old clustering based on the microbial composition at these later timepoints. We built a multi-class XGBoost prediction model using R package XGBoost (v.1.7.9.1) (objective: multi:softprob, default parameters)93. We applied this prediction using a compositional data analysis approach94 by generating all possible log-ratios between bacterial species abundances with >20% prevalence in a leave-one-out cross-validation (LOO-cv) setting. We performed the same analysis using the microbial community from maternal samples at P12, at either P28 or birth and at 3 months postpartum.

Associations between clusters and predictors were performed using logistic regression, using each of the binarized cluster-belonging/-not belonging vectors as an outcome variable and each predictor as the independent variable. We ran a model per predictor and per cluster and controlled for false discovery using the Benjamini–Hochberg FDR procedure. Alpha diversity at 2 weeks, 3 months and 6 months old was associated with clusters using ANOVA, followed by a Tukey post hoc test for pairwise comparisons.

Within and between individual distances in mothers and infants

To calculate the significance of pairwise Aitchison distances between (1) the same individuals; (2) related individuals; and (3) unrelated individuals, we used a permutation-based approach. This involved calculating the t-statistics between the distances of different groups and deriving P values from an empirical null distribution of t-statistics derived from 10,000 permutations of the group labels.

Latent variable analysis assessing the temporal variation in the infant gut microbiome (MEFISTO)

To understand the drivers of temporal variation in the infant microbiome, we performed a FA that assesses the strength of factors (that is, latent variables) through time28 (package MOFA2 (v.1.14.0)), trained using CLR transformed data at SGB abundance level. It therefore permits a longitudinal analysis that estimates latent factors structuring the data across all timepoints and allows inclusion of samples with missing observations.

To establish a stable and appropriate number of factors for MEFISTO, we performed a resampling-based factor stability analysis. MEFISTO models were fitted across a range of factor dimensionalities, from 2 to 12 factors. For each dimensionality, 15 random subsampling splits were generated by selecting 80% of individuals from the full infant microbiome dataset, without replacement. Each subsampled dataset was used to refit the MEFISTO model, yielding estimates of factor scores (Z) and factor weights (W). For each factor dimensionality, factor stability was assessed by computing pairwise Pearson correlations across all combinations of subsampling splits. For the factor weights (W), correlations were calculated after matching corresponding factors across splits based on similarity of their weight vectors, as factor ordering is not identifiable. Factor directions were additionally aligned by sign-flipping to ensure consistent orientation across runs. For factor scores (Z), stability was assessed by first computing the mean factor score across individuals within each subsampling split. After matching and sign-aligning factors across splits, pairwise Pearson correlations were calculated for the mean factor scores across all combinations of subsampling splits at each factor number.

Together, these analyses quantify the stability of inferred factors and their associated weights as a function of the number of estimated factors. Higher correlations indicate greater robustness of the inferred latent structure, thereby informing the selection of an appropriate number of factors for downstream analyses. The three factor solution exhibited near perfect reproducibility of both Z and W across splits, indicating a highly stable latent structure. When adding four or more factors, values of Z and W exhibited lower mean correlations and substantially greater variability across subsampling splits, suggesting overparameterization and reduced reproducibility of the inferred latent structure. We therefore chose a three-factor model. These factors explained 8.27% (factor 1), 4.73% (factor 2) and 3.88% (factor 3) of total variance, respectively, with 19.4% explained by the full model. These factors were moderately correlated (r = 0.40–0.58), consistent with partially overlapping but non-orthogonal dimensions of microbial variation.

We next modelled the inferred factor values as outcomes in penalized mixed-effects regressions to quantify the relative contributions of biological predictors (for example, mode of delivery), temporal variables (for example, time) and technical covariates (for example, sequencing depth) to variation in each latent factor. In selecting predictors from all cross-sectional variables, we included those with a maximum of 15% missingness and excluded those for which the most frequent category accounted for more than 95% of observations. We then retained predictors with a Spearman’s correlation not exceeding 0.5. This resulted in 10 predictors, 3 time variables (the first-, second- and third-degree polynomials of time) and technical variables, such as DNA concentration, sequencing depth and batch number, resulting in 2,428 infant samples with no missing data for those variables. To ensure comparability of penalized coefficients, all predictors were standardized to mean zero and unit variance before model fitting, including continuous and dummy-coded categorical variables. To include the assessment of non-linear time in addition to linear effects of time, we used the base R function poly() to calculate the first three polynomials of time, standardize their unit sum of squares and orthogonalize their values. This results in variables of time that have a mean of zero. Thereafter, all these components of time were scaled to have a s.d. of one, thereby matching the s.d. of all other biological or technical variables.

This set of standardized predictor variables were then used in a penalized regression model that takes into account random effects at the level of the individual. For each factor, a model was fitted using the R package glmmPen (v.1.5.4.8)95, which uses a Monte Carlo expectation–maximization algorithm, with adaptively increasing Monte Carlo sample sizes during optimization, until convergence. The regularization parameter lambda (λ) was selected via the Bayesian information criterion using a penalized likelihood optimization built into glmmPen. Models were fitted using two elastic-net mixing parameters corresponding to pure LASSO (α = 1.0) and a partially ridge-regularized model (α = 0.5) to confirm that our conclusions based on variable selection were not sensitive to the regularization scheme chosen. Results mentioned in the text are for α = 0.5. Extended Data Fig. 4 shows results for both α = 0.5 and α = 1.0.

Associations of infant and maternal predictors with overall gut microbiome composition

In this analysis, we investigated the effects of both cross-sectional predictors, which were static over time (for example, place of delivery), and longitudinal predictors, which varied over time (for example, stool frequency). We first filtered predictors to those present in at least 50 samples at each infant timepoint and at least 100 samples at each maternal timepoint, excluding predictors for which the most frequent category accounted for more than 95% of observations. To test the effect of predictors on the overall gut microbiome composition at each timepoint, we used the adonis2 function from the vegan package in R with 10,000 permutations, using Aitchison distances as described above and correcting for read depth, DNA concentration and batch number for the base model. For infants, this was followed by an extended model including these covariates plus mode of delivery, and a final model incorporating mode of delivery and feeding mode. In all cases, predictors were added to the model with covariates and the setting by=“margin” was used. To test the effect of predictors on the overall composition (termed overall in the main text and figures), considering the temporal patterns of the gut microbiomes of mothers and infants, we used TCAM, a dimensionality reduction method for longitudinal ’omics data analysis29. TCAM is an unsupervised tensor factorization method that enables the decomposition of time-series data, thereby reducing the dimensionality for microbiome studies. For maternal data, three timepoints were selected from pregnancy to three months postpartum (P12, birth and 3 months postpartum). For infants, four timepoints across the first year were selected (1, 3, 6 and 12 months old). Only samples from individuals with all specified timepoints were included in the analysis, as TCAM does not allow for missing data. Gut microbiome profiles were filtered at the SGB level to retain taxa with a relative abundance >0.01% in at least five individuals and were subsequently CLR-transformed. TCAM analysis was conducted according to the protocol described in the original publication29, using the mprod package (v.0.0.5a1)96 in Python. We calculated the Euclidean distance between TCAM components, and associations between the distance matrix and predictors were determined with 10,000 permutations, as described above. These results were then combined with the per-timepoint analysis.

Associations of infant- and maternal-specific predictors and HMOs with the gut microbial species and pathways in mothers and infants

These relationships were investigated both cross-sectionally, where predictors (such as place of delivery) were static over time, and longitudinally, where predictors (such as stool frequency) changed over time. The effects of static predictors on alpha diversity, SGBs relative abundance and pathway abundance were tested with mixed models for repeated measures, using the mmrm package (v.0.3.12)97 in R. For dynamic predictors, we used generalized additive models, using the mgcv package (v.1.9-1)98 in R. In both cases, read depth, sample DNA concentration and batch number were included as covariates in the base model. For infants, this was followed by an extended model including these covariates plus mode of delivery, and a final model incorporating mode of delivery and feeding mode (when these were not the primary predictors of interest). Sample ID was included as a random effect in all models. Numeric variables were inverse-rank transformed before analysis. We included only predictors of interest that were present in more than 200 samples and excluded those for which the most frequent category accounted for more than 95% of observations. Associations of delivery-related predictors with the gut microbiome was performed as above, only in VG infants. To assess associations between breast-milk HMOs and the infant gut microbiome, we used generalized additive models and linked the inverse-rank-transformed HMO concentrations in breast milk at a specific timepoint with the infant gut microbiome at the same timepoint. Associations were corrected for technical variables (DNA concentration, sequencing depth and sample batch), and the analysis was restricted to breastfed individuals. Benjamini–Hochberg correction was used to control for multiple testing in all these analysis, and the number of tests equal to the number of feature–predictor pairs was tested. Results were considered significant at FDR < 0.05.

Association of GBMs with species and predictors

Associations between variables and GBMs were tested on ALR-transformed data using the same modelling framework used for microbial taxa and adjusting for technical covariates, delivery mode and feeding mode when these were not the primary variables of interest.

Association of CAZymes with species and predictors

To assess the influence of bacterial composition on the variation in CAZyme profiles, we tested the CLR-transformed abundances of taxa at the SGB level in mothers and infants using PERMANOVA (adonis) with 10,000 permutations against the Aitchison distance matrix calculated from CLR-transformed CAZyme data. Associations between variables and CAZyme ALR-transformed data were tested using the same modelling framework applied to microbial taxa, adjusting for technical covariates, delivery mode and feeding mode when these were not the primary variables of interest. Hand-annotated substrate specificity per CAZyme was retrieved87. Enrichment analysis was performed using the clusterProfiler::enricher function in R99, using FDR-significant associations between CAZymes and timepoint (early versus late), mode of delivery (VG versus CS) and feeding mode (exclusive breastfeeding versus exclusive formula feeding), tested against a background of all detected CAZymes fulfilling the above-mentioned filter.

Prediction model of maternal gut microbiome and infant eczema

To assess the independent association between maternal gut microbiome alpha diversity and the risk of infant eczema, we first performed a multivariable logistic regression including maternal alpha diversity, smoking history (previously linked to reduced diversity), and family history of allergic disease as covariates. We then evaluated the predictive performance of microbiome-derived features using a leave-one-out cross-validation (LOO-CV) framework with XGBoost classifiers93. Three sets of predictors derived from the maternal gut microbiome at birth were tested: (1) taxonomic features, represented by CLR–transformed SGB abundances of mother at birth (filtered at a 30% prevalence); (2) functional pathways, represented by ALR–transformed MetaCyc pathway abundances at birth (filtered at a 50% prevalence); and (3) alpha diversity, represented by the Shannon diversity index of mother at birth. Predictive performance was assessed using ROC curves and AUC metrics. To identify the most influential microbial features contributing to model predictions, SHAP values were computed for the species-level taxonomic mode100.

Bacterial strain transmission between mothers and infants

Defining strain-sharing

Phylogenetic distance matrices were extracted from maximum likelihood trees. To define phylogenetic distance cut-offs to identify strain-sharing, we followed a previously published approach22. Samples collected within 6 months of each other were assumed to contain the same strain. A genetic distance cut-off is then optimized to separate the distributions of longitudinal samples (assuming they contain the same strain) and unrelated samples (assuming they have different strains). To differentiate between the same and different strains, we established cut-offs for a total of 1,205 species (specific cut-offs per species are shown in Supplementary Table 56). We then investigated strain-sharing between infants at any age (2 weeks and 1, 2, 3, 6, 9 and 12 months) and maternal strains at birth or P28 if birth strains were not available. Per species, interindividual phylogenetic distances were normalized to their maximum distance. We prepared a training distance dataset to define cut-offs for strain-sharing. For this, we defined training samples with the same strain as those from the same subject and taken within 180 days of each other (that is, 6 months). If a single individual had multiple samples taken within 180 days, the shortest time difference was selected. For the training dataset of samples with different strains, we took unrelated samples (not the same individual or family). If multiple timepoints were present between two samples, we randomly selected one.

Using this training dataset, we optimized a Youden cut-off for the distance that better optimized the function ‘sensitivity + specificity – 1’, given that we defined ‘same strain’ as samples collected from the same individual within 180 days, and ‘different strain’ as those collected from unrelated individuals. If the number of training individuals with the same strain was below 50, we used the third quantile of the distribution instead of the Youden cut-off since we believe the Youden cut-offs to be overfitted. If the Youden cut-off was larger than the distance value identified in the fifth quantile, the fifth quantile was used instead as a more conservative metric. Using this approach, we identified cut-offs for 1,205 species. Using the complete distance matrix, we considered two samples to have the same strain of a species if their normalized phylogenetic distance was equal to or lower than the identified cut-off.

Strain-sharing associations with predictors

Per species, we removed all pairs of samples that did not belong to mother–infant pairs (that is, each mother with her own offspring). If we did not identify at least two mother–infant pairs, the species was removed from analysis. We then considered only one mother timepoint, either at birth or P28 (as the closest timepoint to birth), to focus on possible vertical transmission events.

Mother–infant strain sharing was then associated with infant age (in days). For this analysis, we focused only on species with over 20 mother–infant pairs (multiple timepoints of the same infant with the same timepoint from mother were used). We further removed species where there were fewer than 10 mother–infant unique pairs (without counting multiple timepoints), or if there were less than 3 samples with at least 5 timepoints available. This allowed us to test for 49 species. We then ran a generalized mixed-effects model with a logit link (logistic regression), where mother–infant strain sharing was used as the dependent variable, while age in days, mother timepoint (P28 or birth) and infant ID (as a random effect) were used as covariates.

To associate mother–infant strain sharing with predictors, we used species with over 20 mother–infant pairs (multiple timepoints of the same infant and the same timepoint from mother were used). Different predictors had different missing rates, and only those complete in at least ten individuals were used for analysis. However, for categorical predictors, we required at least five strain-sharing events per predictor level. Finally, we performed another logistic generalized linear mixed model using strain sharing as the dependent variable and the predictors, mother timepoint, infant timepoint and infant ID as covariates. All associations were merged, and a Benjamini–Hochberg FDR was estimated.

To associate overall sharing rates (that is, not per species, but rather the percentage of strains shared between a mother and an infant per timepoint), we merged all mother–infant distances (using a single mother timepoint as described above) from all species. Then, per mother–infant pair, we assessed the percentage of species for which the same strain was shared. We included only pairs for which we could assess this rate using at least five species. We then ran linear mixed models using the rate as the dependent variable, which was associated with an independent model per predictor, while also accounting for infant and mother timepoints and infant ID as a random effect. Finally, we estimated an FDR from all the predictor associations.

All associations with the predictors duration of ruptured membranes, duration of delivery and place of delivery were only run in VG infants.

Strain sharing and bacterial abundance

For each species, we examined whether its CLR-transformed abundance in the infant was associated with whether the strain was shared with the mother (at birth or P28 if no birth sample available). As this effect might be different in early timepoints compared to later timepoints, we labelled the data as early if the infant data were from infants aged 2 weeks or 1 , 2 or 3 months and as late if the data were from infants aged 6, 9 or 12 months. To run the model, we required there to be at least five strain-sharing events and not-strain-sharing events in both early and late timepoints. We then ran a linear mixed-effect model in which the (CLR-transformed) bacterial abundance acted as dependent variable, and the independent variables were whether the strain was shared with the mother, whether the timepoint was early or late and an interaction term between strain-sharing and early/late timepoint, in addition to a random effect with sample ID. We ran this model for all species and extracted the effects of strain-sharing and the interaction between strain-sharing and timepoint. FDR was estimated for these two variables among all tests.

We also associated the mother’s bacterial abundance with the likelihood of strain sharing. For that, we took all available mother timepoints and retained only the early infant timepoints. We required at least five strain-sharing events for association. We associated mother–infant strain sharing, as a dependent variable, with the (CLR-transformed) bacterial abundance of mothers, while controlling for mother timepoint, infant timepoint and mother ID as a random effect. FDR were estimated for the effect of abundance in all tested species.

Strain sharing and strain persistence

We defined strain persistence as those strains that were the same within the same infant at an early (2 weeks or 1 month) and late (9 or 12 months) timepoint, choosing the longest distance among the available timepoints. We then matched information about strain persistence with mother–infant strain sharing at either 2 weeks or 1 month. We required at least 20 mother–infant pairs with information about persistence for analysis. We performed a Fisher exact test to estimate whether the odds of being a persistent strain were higher for strains shared between mother and infants (null hypothesis = higher), than those that were not shared. FDRs were estimated.

De novo assembly and binning

Quality-filtered and trimmed reads from each sample (including unmatched reads) were assembled into contigs using MetaSPAdes (v.3.15.5) using the default parameters101. Bacterial binning was performed separately for each metagenome using metaWRAP (v.1.3.2)102. The resulting bins were dereplicated using skDER (v.1.2.7)103 with the dynamic approach and the following parameters: –percent-identity-cutoff 98.0 and –aligned-fraction-cutoff 90.0.

Functional enrichment of transmitted bacterial strains

High-quality bacterial MAGs (completeness ≥ 90%, contamination < 5%), dereplicated at the subspecies level (98% ANI), were assigned SGB taxonomy using PhyloPhlAn’s phylophlan_assign_sgbs with the mpa_vJan21 MetaPhlAn database, based on MASH distance (<0.05). In total, we identified 645 MAGs that could be assigned to 112 of these SGBs. Metagenomic reads were mapped to selected MAGs using Bowtie2 (v.2.5.1), and mapping files were processed with inStrain (v.1.9.0)104 using the profile module (minimum mapQ score = 0; insert size = 160). inStrain ‘compare’ was used to estimate genome similarity across samples by comparing profiles with a minimum genome breadth ≥ 0.5. Only regions with ≥5× coverage were included in the comparison. Sample pairs with <50% comparable regions of the genome were excluded. Strain sharing was assessed based on population-level ANI (popANI), with bacterial strains considered shared between mother and infant if they exhibited ≥99.999% popANI across comparable regions. Protein sequences in MAGs were predicted using prodigal (v.2.6.3)105 and were functionally annotated using eggNOG-mapper (v2.1.12)106 with the ‘–m hmmer -d 2 –evalue 1e-05’ parameters to retrieve COG annotations. Functional enrichment was assessed by testing whether COGs (encoded in ≥20 MAGs) were transmitted across more unique mother–infant pairs than expected by chance, using a permutation-based test with 10,000 iterations. An FDR < 0.05 was considered significant.

Association of infant age and maternal and infant traits to SGB phylogeny

To associate time and maternal and infant traits with SGB phylogeny, we used multivariate mixed model distance matrix regression. RAxML phylogenetic trees were transformed to distance matrices by their branch lengths using the cophenetic.phylo function from R package ape (v.5.8-1)107. Next, using the MDMR package (v.0.5.2)108, we performed mixed model distance matrix regression treating SGB distance as the outcome. To associate phylogenies to time, we included timepoint (factor) and technical covariates (DNA concentration and sequencing depth) as fixed effects and individual ID as random effects. Associations that passed a FDR (Benjamini–Hochberg) cut-off of 5% were considered significant.

$${\rm{Phylogenetic\; distance\; \sim \; covariates\; +\; Timepoint\; +\; (1|Individual\; ID)}}$$

To associate species phylogenies to maternal and infant predictors, we executed the models with timepoint adjusted as fixed effect:

$$\begin{array}{c}{\rm{Phylogenetic\; distance}} \sim {\rm{covariates}}+{\rm{Timepoint}}+{\rm{Predictor}}\\ \,+(1|{\rm{Individual\; ID}})\end{array}$$

Quantitative predictors were normalized using inverse rank transformation. As the sample size varied across the trees tested (due to differences in species prevalence), we included only predictors of interest that were present in more than 100 samples and for which the most frequent category comprised no more than 75% of observations.

The phylogeny–predictor associations with analytical MDMR P values that passed FDR (Benjamini–Hochberg) control of 5% were additionally validated by performing 20,000 permutations and calculating empirical P values. The associations were considered significant if both the analytical and empirical P value reached the 5% FDR threshold.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.



Source link

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use
Add a comment Add a comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Previous Post
Ultrafast temporal engineering of optical properties using terahertz light

Ultrafast temporal engineering of optical properties using terahertz light

Next Post
No single rule explains how climate affects carbon storage in tropical forests

No single rule explains how climate affects carbon storage in tropical forests

Advertisement