Bacterial strains
All strains for the experiments were grown in lysogeny broth (LB) Lennox and the required antibiotics and incubated at 37 °C with shaking. E. coli K-12 (BW25113)72 was used for most biological analyses, and the BL21-AI or NEB5α strains were used for protein purification and cloning, respectively. The use of the ECOREF53 strain collection is described in detail in later sections. When culturing strains transformed with plasmids, the following antibiotic selection was used: 100 μg ml−1 spectinomycin, 25 µg ml−1 chloramphenicol and 100 μg ml−1 kanamycin.
Phages
Phage T7 WT was acquired from DSMZ, T7Δ0.7::cmk (referred to as Δ0.7 throughout the manuscript) was a gift from U. Qimron73. Our phage stocks were checked for the presence of the correct T7 genotype with the primers TAGCCAACACACTGAACGCT and CACCCGCAGATGACCTGTAA for T7 WT and TAGCCAACACACTGAACGCT and GCCAAAAGGCGCTCAAAGTT for T7Δ0.7::cmk. Moreover, we confirmed that the known deletion mutant H1 (ref. 74) did not occur in our stocks (using the primers TAGCCAACACACTGAACGCT, CACCCGCAGATGACCTGTAA and GCCAAAAGGCGCTCAAAGTT). Two mutants of T7, the catalytic dead G76F35 and a truncated kinase mutant ΔSO (kinase domain only, residues 1–242), were generated using in vitro methods for genome engineering and cell-free phage rebooting using proprietary methods developed by Invitris. Phages from the BASEL collection59 (Bas64, Bas65, Bas66(T3K12), Bas67 and Bas68) were a gift from A. Harms. Note that the infection curves of our Bas67 stock behaved significantly differently to the other BASEL Autographiviridae, so we excluded this phage from further analyses. Phage P1vir was single-plaque purified from a Typas laboratory stock. Annotated genomes were downloaded from the NCBI using the GenBank accessions as reported in ref. 59 and NC_047864.1 (reference genome used to represent T3/Bas66(T3K12)) and NC_005856.1 (reference genome used to represent P1vir).
Phages were propagated in the E. coli K-12 (BW25113) strain and cultured in LB, 5 mM CaCl2 and 10 mM MgCl2 at 37 °C with shaking. Lysates of overnight cultures infected with phage stocks were prepared by adding chloroform, vortexing and centrifuging (4,000g, 4 °C, 5 min) to remove debris. The lysate was transferred to fresh tubes, additional chloroform was added, and the samples were vortexed and centrifuged again. Clarified lysates were stored at 4 °C. Lysate titres were determined in duplicates using the double-agar overlay method75.
Plasmids
For the study of recombinant T7K, the regions encompassing the N-terminal PK domain (residues 1–242) and the C-terminal SO domain of T7K (residues 243–359) were PCR-amplified from T7 gDNA using Q5 High-Fidelity DNA Polymerase (NEB), with primers designed to introduce a 6×His-TEV N-terminal fusion and carrying overhangs to plasmid pET28a.
Almost all plasmids used in the T7K screen (Supplementary Table 5) are in the pTU175 backbone, which carries the low-copy-number origin of replication pSC101, except for the Borvo and Retron-Eco8 defence systems, which were obtained from the lab of Rotem Sorek and kept in the pSG1-vector, expressed from their native promoter and with a p15a origin of replication48,52, and CmdTAC (PD-T4-9) defence system, which was obtained from the laboratory of M. Laub, and was kept on its original plasmid pCD1–PD-T4-9 under its native promoter76. Plasmids containing Retron-Eco1, Retron-Eco9 and Retron-Sen2 have been described previously33. For Retron-Eco3, Retron-Eco6 and Retron-Eco10, the retrons were amplified with their native promoters (500 bp upstream of the retron start) from strains from the ECOREF E. coli natural isolate library53 or from strain SC402 (Retron-Eco3), and cloned into pTU175. For the toxin–antitoxin system TacAT77, the protein-coding sequence was amplified from Salmonella typhimurium LT2 and cloned into the pTU175 backbone under the arabinose-inducible pBAD promoter. The HEPN-MNT78 and the ApeA49 defence systems were cloned into the pTU175 vector under the arabinose-inducible pBAD promoter (gift from I. Songailiene, V. Šikšnys laboratory). The DarTG1 defence system was cloned into the pTU175 vector under the arabinose-inducible pBAD promoter by M. LeRoux79.
Phosphomimetic mutant constructs of RcaT (in Retron-Eco9) and DarT (in DarTG1) were generated using the Q5 site-directed mutagenesis kit (NEB), using the primers described in Supplementary Table 6. Plasmid edits were confirmed by Sanger sequencing and PacBio whole-plasmid sequencing.
Infection sampling (single timepoint)
To measure proteome and phosphoproteome changes in several phage infections (WT T7, T7 mutants, BASEL phages) of E. coli K-12 (BW25113), a single colony was picked and grown overnight in LB, 10 mM MgCl2 and 5 mM CaCl2 and, the next morning, was back diluted 1:1,000 into fresh LB, 10 mM MgCl2 and 5 mM CaCl2 (300 ml). The bacteria were cultured at 37 °C with shaking, and their optical density at 600 nm (OD600) was monitored over time until an OD600 of around 0.5 (1 cm pathlength). Phage stocks were added to a final MOI of 5 and infections were allowed to proceed at 37 °C with shaking. Uninfected samples were immediately placed on ice. To collect at the indicated timepoints, the flasks were immediately transferred to a pre-cooled centrifuge and spun (4,000g, 5 min, 4 °C). For each sample, the supernatant was removed and the cell pellet was resuspended in 1 ml ice-cold PBS, and then centrifuged once again (10,000g, 1 min, 4 °C). The supernatant was aspirated and the pellet was flash-frozen in liquid nitrogen. In total, three biological replicates were performed for each phage infection state (three independent colonies prepared on different days). The samples were then prepared for phosphoenrichment (see below), and matched protein expression measurements of the unenriched peptide digests also collected using data-independent analysis without further offline peptide fractionation (see below).
Infection time-course sampling
To measure proteome and phosphoproteome changes in WT T7 and Δ0.7 phage infections of E. coli K-12 (BW25113), samples were prepared as above but the culture was then split into six independent conical flasks and phage stock (either WT or Δ0.7, or no stock as the uninfected control) added to a final MOI of 5 and infections were allowed to proceed at 37 °C with shaking. Uninfected samples were immediately placed on ice, or for infections collected at 1, 5 or 10 min after infection for WT T7, or at 5 and 10 min for Δ0.7, the flasks were immediately transferred to a precooled centrifuge and spun (4,000g, 5 min, 4 °C). In total, three biological replicates were performed (time courses from three independent colonies prepared on different days).
To compare phosphoproteome changes in WT T7 and ΔSO phage infections, samples from each phage were collected as described above at 10 min after infection (to measure the effect of the loss of the SO domain at saturation), or at 2, 3 or 4 min after infection (to measure the effect before saturation), with three biological replicates performed for both experiments (time courses from three independent colonies prepared on different days for each experiment). To increase temporal resolution for the pre-saturation experiment, we collected the cell suspension directly into Falcon tubes prefilled with ice cold PBS.
MS sample preparation (time course)
The lysis buffer was composed of 4 M guanidinium isothiocyanate, 50 mM HEPES, 5 mM TCEP, 20 mM chloroacetamide, 1% N-lauroylsarcosine, 5% isoamyl alcohol and 40% acetonitrile, with the pH adjusted to 8.5 using 10 M NaOH. For cell lysis, a buffer volume of approximately five times the cell pellet volume was added. The samples were homogenized by pipetting, incubated on a shaker at room temperature for 1 h, and centrifuged at 16,000g for 10 min to remove cell debris and nucleic acid aggregates. The protein concentration was determined by tryptophan fluorescence, as previously described80. The samples were transferred to MultiScreenHTS-HV 96-well filter plates with 0.45 µm PVDF membranes (Merck Millipore), and ice-cold acetonitrile was added to each well to induce protein precipitation, achieving a final concentration of 80% acetonitrile. After a 10 min incubation, the samples were centrifuged, and the supernatant was discarded. The protein precipitates were then washed twice with 200 µl of 80% acetonitrile and twice with 200 µl of 70% ethanol, with each wash followed by centrifugation at 1,000g for 2 min.
Next, a digestion buffer containing 50 mM triethylammonium bicarbonate buffer and trypsin (TPCK-treated, Thermo Fisher Scientific) was added to the protein precipitates. The trypsin-to-protein ratio was set to 1:25 (w/w), with a maximum final protein concentration of 10 µg µl−1. Tryptic digestion was performed overnight at room temperature with mild shaking (600 rpm). After digestion, the samples were dried using a vacuum concentrator.
For the parallel proteomics measurements, the full digest was TMT labelled (see below) and then subjected to offline high-pH fractionation (see below). For phosphoproteomics measurements phosphopeptide enrichment was performed (see below).
Phosphopeptide enrichment
For phosphoproteomics studies (single infection timepoint for several phages, time courses comparing WT T7 with Δ0.7 or with ΔSO infections, Retron-Eco9 and DarTG1 infections), phosphopeptide enrichment was performed81. Lyophilized peptides were resuspended in a loading and washing buffer composed of 80% acetonitrile and 0.07% TFA, sonicated and centrifuged at 16,000g. Phosphopeptide enrichment was carried out on the supernatant as described previously82 using the KingFisher Apex robot (Thermo Fisher Scientific) with 25 µl of Fe-NTA MagBeads (PureCube) per sample. After five washes with buffer A, phosphopeptides were eluted with 100 µl of 0.2% diethylamine in 50% acetonitrile, followed by lyophilisation.
The enriched phosphoproteomic time-course samples were TMT labelled (see below) and then subjected to offline porous graphitic carbon (PGC) fractionation. The enriched single timepoint infections and Retron-Eco9 and DarTG1 infection samples were not TMT-multiplexed and were subjected to MS analysis without further peptide fractionation.
MS sample preparation (defence systems)
For phosphoproteomics studies, E. coli K-12 were transformed with plasmids encoding Retron-Eco9 (native expression) or DarTG1 (arabinose-inducible). Overnight cultures were started in LB, 10 mM MgCl2, 5 mM CaCl2, 100 μg ml−1 spectinomycin, and then from these a 50 ml day culture (1:2,000 back dilution) was started the next morning in fresh LB, 10 mM MgCl2, 5 mM CaCl2, 100 μg ml−1 spectinomycin. DarTG1 culture medium had 0.2% (w/v) arabinose for the induction of DarTG1 expression. Cells were grown until an OD600 of around 0.05 (Retron-Eco9) or about 0.6 (DarTG1) (pathlength = 1 cm), infected with WT T7 at an MOI of 5 and then collected after 5 min of infection, as described above in the ‘MS sample preparation (time course)’ and ‘Phosphopeptide enrichment’ sections. Two independent biological replicates were performed.
For experiments assessing phosphomutant protein stability (through abundance), cells transformed with either Retron-Eco9 or DarTG1 constructs (or their corresponding empty plasmids) were cultured as described above, but collected (without phage infection) at an OD600 of around 0.05 (Retron-Eco9) or 2 (DarTG1). Cells were collected by centrifugation (5 min, 4,000g, 4 °C), washed once with ice-cold PBS, centrifuged again and then resuspended in 5 pellet volumes of 2% (w/v) SDS. The samples were boiled for 5 min, and then subjected to five rounds of freeze–thaw before being processed for MS analysis (sample cleanup, proteolytic digestion, TMT-multiplexing, high pH fractionation; see the sections below). Three independent biological replicates were performed.
MS sample clean-up and protein digestion
MS sample preparation and measurements were performed as described previously83. Protein digestion was performed using a modified SP3 protocol84. The samples were diluted to a final volume of 20 μl with 0.5% SDS and mixed with a paramagnetic bead slurry (10 μg beads (Sera-Mag Speed beads, Thermo Fisher Scientific) in 40 μl ethanol). The mixture was incubated at room temperature with shaking for 15 min. The beads bound to the proteins were washed 4 times with 70% ethanol. Proteins on beads were reduced, alkylated and digested using 0.2 μg trypsin, 0.2 μg LysC, 1.7 mM TCEP and 5 mM chloroacetamide in 100 mM HEPES, pH 8. After overnight incubation, the peptides were eluted from the beads, dried under vacuum, reconstituted in 10 μl of water and labelled with TMTpro, 18-plex reagents and dissolved in acetonitrile at 1:15 (peptide:TMT weight ratio) for 1 h at room temperature. The labelling reaction was quenched with 4 μl of 5% hydroxylamine and the conditions were pooled together. The pooled sample was desalted with solid-phase extraction after acidification with 0.1% formic acid. The samples were loaded onto the Waters OASIS HLB μelution plate (30 μm), washed twice with 0.05% formic acid and finally eluted in 100 μl of 80% acetonitrile containing 0.05% formic acid. The desalted peptides were dried under vacuum and dissolved for analysis using liquid chromatography coupled with MS (LC–MS).
Expression and purification (6×His-TEV-PK)
The plasmid pET28a-6×His-TEV-PK was transformed into E. coli BL21-AI. Cells were grown at 37 °C in 1.2 l of ZYM-5052, supplemented with 100 μg ml−1 kanamycin (Sigma-Aldrich) and 0.05% arabinose (Sigma-Aldrich), and cultures were shifted to 30 °C when the OD600 reached around 0.8–0.9. Protein overproduction was carried out for a further 18 h. Cells were collected in the Beckman-Coulter JLA 8.1000 rotor (5,000 rpm, 30 min, 4 °C) and pellets were resuspended in 45 ml of binding buffer (20 mM Tris pH 7.5, 500 mM NaCl, 5 mM imidazole) supplemented with EDTA-free protease inhibitor (Roche), 100 μg ml−1 DNase I (Sigma-Aldrich) and 0.5 mg ml−1 lysozyme (Sigma-Aldrich). Cells were lysed in a C3 Emulsiflex high-pressure homogenizer (Avestin) and insoluble material was pelleted (80,000g, 20 min, 4 °C). The supernatant was passed through a 0.2 μm PVDF filter (Merck Millipore) and applied to 1 ml of pre-equilibrated Ni-NTA agarose beads (Qiagen) on a gravity-flow column. Beads were washed with 10 ml of binding buffer, then 10 ml of wash buffer (20 mM Tris pH 7.5, 500 mM NaCl, 60 mM imidazole) and the protein eluted in 5 ml of elution buffer (20 mM Tris pH 7.5, 500 mM NaCl, 1 M imidazole), collecting 500 μl aliquots. The fractions with the highest purity and yield were pooled and dialysed twice against 2 l of dialysis buffer (20 mM Tris pH 7.5, 500 mM NaCl). The protein was stored at −80 °C in 30% glycerol.
Expression and purification (6×His-TEV-SO)
The plasmid pET28a-6×His-TEV-SO was transformed into E. coli BL21-AI. Cells were grown at 37 °C in 1.2 l of ZYM-5052 (ref. 85), supplemented with 100 μg ml−1 kanamycin (Sigma-Aldrich) and 0.05% arabinose (Sigma-Aldrich), and cultures were shifted to 30 °C when the OD600 reached about 0.8–0.9 (pathlength = 1 cm). Protein overproduction was carried out for a further 18 h. Cells were collected in the Beckman-Coulter JLA 8.1000 rotor (5,000 rpm, 30 min, 4 °C) and the pellets resuspended in 45 ml of lysis buffer (50 mM Tris pH 7.5, 5 mM EDTA) supplemented with EDTA-free protease inhibitor (Roche), 100 μg ml−1 DNase I (Sigma-Aldrich) and 0.5 mg ml−1 lysozyme (Sigma-Aldrich). Cells were lysed in a C3 Emulsiflex high-pressure homogenizer (Avestin) and insoluble material pelleted (80,000g, 20 min, 4 °C). The pellet was resuspended in 30 ml of membrane solubilization buffer (50 mM Tris pH 7.5, 2% Triton X-100) and samples were pelleted again (80,000g, 20 min, 4 °C). The pellet was then resuspended in high salt buffer (50 mM Tris pH 7.5, 1 M NaCl, 2% Triton X-100) and ultracentrifuged again (80,000g, 20 min, 4 °C). Inclusion bodies were resuspended in high-urea buffer (50 mM Tris pH 8.0, 7.6 M urea) and let solubilize at 4 °C for 18 h with gentle shaking. Insoluble material was removed by ultracentrifugation (80,000g, 20 min, 4 °C) and urea-solubilized protein was stored at −80 °C.
T7K protein solubility (western blotting)
Lysates from cells expressing either 6×His-TEV-PK or 6×His-TEV-SO were mixed with SDS–PAGE loading buffer either directly after lysis (whole cell) or after removal of insoluble material by centrifugation at 16,000g and 4 °C for 45 min (soluble). The samples were boiled at 100 °C for 10 min, then loaded on 15% polyacrylamide for SDS–PAGE. Proteins were transferred to PVDF and membranes blocked in Tris-buffered saline buffer containing 0.5% Tween-20 (TBST) and 5% skimmed milk for 1 h at room temperature. Membranes were incubated with anti-His (Invitrogen, MA1-21315) at 1:2,000 overnight at 4 °C, then washed three times in TBST before incubation with goat anti-mouse HRP (Cytiva, NA931) at 1:10,000 for 1 h at room temperature. Membranes were washed again three times in TBST, then imaged on a iBright1500 system (Thermo Fisher Scientific).
In vitro phosphorylation assay
To remove autoinhibitory phosphosites on T7K and enable the study of its activity in vitro, dephosphorylation of purified T7K kinase domain was performed. The reaction was performed in a total volume of 4.5 ml. Purified 6×His-TEV-PK was supplemented with 10 mM MnCl2 and 4,000 U of lambda protein phosphatase (NEB) and incubated at 30 °C. After 3 h, another 4,000 U of lambda protein phosphatase were added, and the sample was incubated at 30 °C for another 3 h. To remove lambda protein phosphatase, the mixture was applied to 500 μl of pre-equilibrated Ni-NTA agarose beads (Qiagen) on a gravity flow column. The beads were washed with 20 ml of binding buffer (20 mM Tris pH 7.5, 500 mM NaCl, 5 mM imidazole), then 6×His-TEV-PK eluted in 5 ml of elution buffer (20 mM Tris pH 7.5, 500 mM NaCl, 1 M imidazole), collecting 500 μl aliquots. The fractions with the highest purity and yield were pooled and dialysed twice against 2 l of dialysis buffer (20 mM Tris pH 7.5, 500 mM NaCl, 1 mM DTT, 1 mM EDTA). The protein was stored at −80 °C in 30% glycerol at a final concentration of 0.5 mg ml−1.
For the in vitro phosphorylation assay, 500 μg of E. coli K-12 tryptic digest per replicate were resuspended in a buffer containing 50 mM HEPES pH 7.5, 100 mM NaCl, 5 mM Mg-ATP and 1 mM TCEP. Subsequently, 5 μg of the purified (and dephosphorylated) kinase domain was added, and the reaction was incubated at 37 °C for 4 h. The peptides were next lyophilized before phosphopeptide enrichment and label-free LC–MS/MS. In total, three biological replicates were performed.
Phosphorylation stoichiometry estimation
Tryptic peptides (40 µg for each of the three replicates) corresponding to either E. coli infected with T7 WT or Δ0.7 (at T = 10 min after infection, three biological replicates) were subjected to alkaline phosphatase treatment (5 U of FastAP per sample; Thermo Fisher Scientific) in a volume of 50 µl of phosphatase buffer, at 37 °C for 20 h. As a control, the same amount of tryptic peptides was incubated in the same conditions without phosphatase. After incubation, the samples were lyophilized and desalted using t-C18 Sep-Pak columns (Waters) before TMT labelling and offline PGC fractionation.
6×His-TEV-SO in vitro refolding assay
gDNA was isolated from 50 ml of E. coli BW25113 overnight culture using the DNeasy Blood & Tissue Kit (Qiagen) and concentrated to about 0.8 mg ml−1 in water in a vacuum concentrator. Total E. coli RNA was obtained from Invitrogen. Refolding reactions were assembled as follows: gDNA (80 μg) or an equal volume of water (no DNA control) were mixed with 40 μg of urea-solubilized 6×His-TEV-SO in 25 mM Tris pH 8.0, 3.8 M urea, in a total volume of 200 μl. For refolding experiments in the presence of RNA, total RNA (40 µg) or an equal volume of water (no RNA control) were mixed with 20 μg of urea-solubilized 6×His-TEV-SO in 25 mM Tris pH 8.0, 4 M urea, 100 U ml−1 SUPERase RNase inhibitor (Invitrogen) in a total volume of 200 μl. Mixtures were transferred to 6–8 kDa molecular weight cut-off D-Tube Dialyzer Mini tubes (Merck-Millipore) and dialysed for 18 h at 4 °C with gentle stirring in refolding buffer (25 mM Bis-Tris pH 6.6, 500 mM NaCl, 1 mM MgCl2). Dialysed samples were collected into clean 1.5 ml microcentrifuge tubes and samples containing RNA supplemented with another 100 U ml−1 RNase inhibitor. Aliquots (15 μl) of total material (input fraction) were collected after dialysis, then the rest of the sample was centrifuged (16,000g, 45 min, 4 °C) and the supernatants (soluble fraction) were collected. Benzonase-treated controls were prepared as follows: 50 μl aliquots of soluble fraction from samples containing gDNA were supplemented with 125 U of benzonase (Merck-Millipore) and 10 mM MgCl2 and incubated at 37 °C for 4 h, then insoluble material was removed by centrifugation (16,000g, 45 min, 4 °C) and the supernatants were collected. In total, samples across three biological replicates were collected.
To assess integrity of RNA after overnight dialysis in refolding experiments, 20 µl of input and soluble fractions from samples containing RNA were mixed with 6× Purple loading dye (New England Biolabs), denatured at 70 °C for 3 min and immediately placed on ice, then resolved on 1.5% agarose in TAE buffer at 80 V. Aliquots of original RNA stock samples were included as integrity controls. RNA was imaged under UV light after ethidium bromide staining.
MS sample preparation of all fractions (input, soluble, benzonase-treated controls), including TMT-multiplexing was prepared as described above (Sample clean-up and proteolytic digestion). As there was some evidence of remnant protease from the genomic extraction kit, for protein digestion, the ‘nonspecific’ setting was used as protease with an allowance of maximum 2 missed cleavages requiring a minimum peptide length of 7 amino acids.
TMT labelling
For the full proteome measurements, 10 µg of peptides per condition was labelled while, for the phosphoproteome, the totality of enriched phosphopeptides was used. Peptides were resuspended in 10 µl of 100 mM HEPES (pH 8.5), and 4 µl of TMTPro reagent at a concentration of 20 µg µl−1 in acetonitrile was added. The labelling reaction was allowed to proceed for 1 h at room temperature, and was then quenched by the addition of 5 µl of 5% hydroxylamine for 15 min. Labelled peptides belonging to the same experiment were subsequently pooled and lyophilized.
Before fractionation, the phosphopeptides were resuspended in 50 µl of 10% TFA and desalted using in-house C18 stage tips86 packed with 1 mg of ReproSil-Pur 120 C18-AQ 5 µm material (Dr. Maisch) above a C18 resin plug (AttractSPE Disks Bio, C18, Affinisep); in the case of the full proteome samples, the peptides were desalted using the Sep-Pak t-C18 column (Waters).
PGC-LC offline peptide fractionation
The samples were resuspended in 18 µl of buffer A (0.05% TFA in MS-grade water supplemented with 2% acetonitrile), with an injection volume set to 16 µl. Peptide separation was conducted using a Hypercarb column (100 mm length, 1.0 mm inner diameter, 3 µm particle size, Thermo Fisher Scientific) at 50 °C, with a flow rate of 75 µl min−1, on an Ultimate 3000 Liquid Chromatography system (Thermo Fisher Scientific). After a 1 min post-injection delay, a linear gradient separation was applied, increasing from 13% buffer B (0.05% TFA in acetonitrile) to 42% buffer B over 47 and 95 min for the phosphoproteome and full proteome samples, respectively. This was followed by an increase to 80% buffer B within 5 min. The column was then washed with 80% buffer B for 5 min before re-equilibration with 100% buffer A for another 5 min. Fractions were collected from 4.5 to 52.5 min and from 4.5 to 100.5 min at 2-min intervals, yielding 24 fractions and 48 fractions for the phosphoproteome and full proteome samples, respectively. These were then pooled into 12 or 24 fractions by combining each fraction with its corresponding n + 12 and n + 24 fraction for the phosphoproteome and full-proteome samples, respectively. After pooling, the fractions were dried using a vacuum concentrator prior to LC–MS/MS analysis.
High-pH offline peptide fractionation
The samples were fractionated using a reversed-phase C18 system running under high-pH conditions, as described previously83. In brief, this consisted of an 85 min gradient (mobile phase A: 20 mM ammonium formate (pH 10); mobile phase B: acetonitrile) at a flow rate of 0.1 ml min−1, starting at 0% B, followed by a linear increase to 35% B from 2 min to 60 min, with a subsequent increase to 85% B up to 62 min, held at 85% B until 68 min, followed by a linear decrease to 0% B up to 70 min, and finally held at this level until the end of the run. Fractions were collected every 2 min from 12 min to 70 min, and fractions were concatenated to make a total of 6 fractions (Retron-Eco9, DarTG1 expression analysis) or 24 fractions (phage infection time-course expression analysis) by combining each fraction with its corresponding n + 6 and n + 24 fraction.
Data-dependent LC–MS/MS
Before injection, all of the samples were resuspended in a loading buffer composed of 1% TFA, 50 mM citric acid and 2% acetonitrile in MS-grade water. LC separation was performed on the UltiMate 3000 RSLCnano system (Thermo Fisher Scientific). Peptides were initially trapped on a cartridge (Precolumn: C18 PepMap 100, 5 μm, 300 μm inner diameter × 5 mm, 100 Å) before being separated on an analytical column (Waters nanoEase HSS C18 T3, 75 μm × 25 cm, 1.8 μm, 100 Å). Solvent A was composed of 0.1% formic acid with 3% DMSO in LC–MS-grade water, while solvent B contained 0.1% formic acid with 3% DMSO in LC–MS-grade acetonitrile. Peptides were loaded onto the trapping cartridge at 30 μl min−1 with solvent A for 5 min, then eluted at a constant flow rate of 300 nl min−1 using a linear gradient of buffer B, followed by an increase to 40% buffer B, a wash at 80% buffer B for 4 min and re-equilibration to the initial conditions. The linear gradient corresponded to an increase from 5% to 25% B in the case of label-free peptides, and from 7% to 27% B in the case of TMT-labelled peptides.
The LC system was coupled to a Fusion Lumos Tribrid, an Exploris 480 mass or a Q-Exactive Plus mass spectrometer (Thermo Fisher Scientific), operated in positive-ion mode. The mass spectrometers were operated in data-dependent acquisition mode with a maximum duty cycle time of 3 s or a top 20 method, selecting precursors with charge states 2–7 and a minimum intensity of 2 × 105 for subsequent HCD fragmentation. Peptide isolation was performed using the quadrupole with 0.7 m/z or 1.4 m/z isolation windows in the case of TMT-labelled and label-free samples, respectively. MS/MS spectra were acquired in profile mode using the Orbitrap, with a maximum injection time of 100 ms and an AGC target of 1 × 105 charges.
Data-independent LC–MS/MS
For protein expression analysis of BASEL phage infection samples, peptides were resuspended in a loading buffer composed of 1% TFA and 2% acetonitrile in MS-grade water and then separated using the Vanquish Neo UHPLC system (Thermo Fisher Scientific) operated in trap-and-elute mode. The LC systems was equipped with a trapping cartridge (Precolumn; PepMap Neo C18, 5 μm, 300 μm inner diameter × 5 mm, 100 Å) and an analytical column (Ionopticks AUR3-25075C18-XT, 25 cm × 75 μm inner diameter, 1.7 μm C18). Solvent A was 0.1% formic acid supplemented with 3% DMSO in LC–MS-grade water and solvent B was 0.1% formic acid supplemented with 3% DMSO in LC–MS-grade acetonitrile. Peptides were loaded onto the trapping cartridge and separated on the analytical column using a linear gradient of 5–26% B (flow rate of 300 nl min−1) for 29.7 min, followed by an increase to 40% B within 3 min before washing at 85% B for 5 min and re-equilibration to initial conditions (total MS acquisition time of 40 min). The LC system was coupled to an Orbitrap Astral mass spectrometer (Thermo Fisher Scientific) operated in data-independent acquisition mode. The instrument was operated in positive-ion mode with a spray voltage of 1.8 kV and a capillary temperature of 280 °C. Full-scan MS spectra were acquired in the Orbitrap at a resolution of 240,000 with a mass range of 430–680 m/z, an AGC target of 5 × 106 charges and a maximum injection time of 5 ms. DIA spectra were acquired in the Astral mass analyser with 2 m/z windows between 430 and 680 m/z. MS2 scan range was set to 150–2,000 m/z, the normalized collision energy to 25 and the default charge state to 2+. The normalized AGC target was set to 500% and the maximum injection time to 5 ms.
LC–MS/MS raw data processing
For all proteomic analyses, a combined protein sequence database was constructed consisting of the Swissprot E. coli strain K-12 database (4,401 entries, UP000000625) supplemented with a common protein contaminants database (118 entries). For analyses involving phage T7 infections, this was further supplemented with the Swissprot phage T7 proteome (UP000000840, 57 entries). For analyses involving the mutant phages, the FASTA including the full length T7K sequence was used. For analyses involving E. coli transformed with either defence-system-containing plasmid, the database was also supplemented with the proteins encoded on each plasmid—DarTG1 (5 entries) and Retron-Eco9 (5 entries). For the SO-domain solubility analysis, the Swissprot E. coli strain K-12 database was supplemented with contaminants and additionally supplemented with the sequence for the recombinant T7K SO-domain protein. For experiments involving BASEL phages (Bas64, Bas65, Bas68 and P1vir), the E. coli database was supplemented with protein FASTAs downloaded from the NCBI (GenBank: MZ501110.1, MZ501081.1, MZ501078.1, MZ501055.1 and NC_005856.1, respectively). The proteome produced by the T3 reference genome GenBank NC_047864.1 was used for Bas66(T3K12).
For data-dependent analyses, raw files were converted to mzmL files using MSConvert (v3.0.24260-9bf2070) from Proteowizard87, using peak picking and keeping the 1,000 most intense peaks per spectrum. Files were then searched using MSFragger (v.4.0)88 in Fragpipe (v.21) using the appropriate database for the given phage infection (see above). The default MSFragger phospho (for label-free) or TMT16-phospho (for proteomics and phosphoproteomics) workflow was used, with a few modifications: oxidation on methionine (maximum 2 occurrences), phosphorylation on Ser/Thr/Tyr (maximum three occurrences) and, for TMT analyses, peptide N-terminal TMT16 labelling (maximum 1 occurrence) set as variable modifications. A total of up to five variable modifications was allowed per peptide. Lysine TMT16 labelling (for TMT analyses) and cysteine carbamidomethylation were set as fixed modifications. Percolator (v.3.6.4)89 was used for PSM validation and PTMProphet (TPP v.6.3.2)90 was used to determine site localization. An FDR cut-off of 1% was used. For the TMT quantification, Philosopher (v.5.1.0)91 was used to extract MS1 and TMT intensities.
For data-independent analyses, raw files were directly inputted into DIA-NN92 (v.2.1.0). Samples were processed based on the identity of the infecting phage (uninfected, WT T7 or BASEL phage). Spectral libraries were generated directly from the inputted FASTA using the deep learning-based spectra, retention time and ion mobility prediction mode, using trypsin/P protease specificity. DIA-NN was run with defaults, except using a precursor and fragment ion m/z range of 430–680 and 200–1,800, respectively, up to two missed cleavages allowed. Furthermore, cysteine carbamidomethylation (+57.021464 Da, UniMod:4) was set as a fixed modification, and methionine oxidation (+15.994915 Da, UniMod:35) as a variable modification, with up to two variable modifications per peptide. Quantification was performed using Legacy (direct) mode summarizing at the protein group level and using the –ids-to-names mode, with the match-between-runs feature. The protein group quantification matrices generated by DIA-NN were used for downstream statistical analyses.
MS data analysis and statistics (general)
For the summarization of information at the phosphosite level, the highest localization score for each site across redundant peptide spectrum matches (PSMs) was considered. Class I phosphosites were those that were uniquely mapping (only one protein in the FASTA database) with a PTMProphet90 phosphorylation localization probability of ≥0.75. The summarization of quantitative information at the (phospho)peptide levels (for phosphoproteomics and stoichiometry experiments) was performed by summing the TMT intensities of all the PSMs assigned to that specific (phospho)peptide.
For TMT analyses, the PSM tables produced by FragPipe were filtered to retain PSMs with a purity value of >0.5 with full quantification (the signal in all used TMT channels). For all experiments, contaminant proteins were filtered out. For phosphoproteomics experiments, all phospho-modified peptides were considered; for the stoichiometry experiment, only unique peptides were considered; for protein expression analysis, only proteins with at least one unique peptide were considered. Differential expression analysis was performed using the summed TMT reporter ion intensities (summarized at the modified peptide level) for phosphoproteomics and stoichiometry experiments, or using the protein intensities from MSFragger output for other protein expression experiments. Vsn-normalization93 and differential expression analysis were used as implemented in the limma package (v.3.58.1)94 in R (v.4.3.3). For the time courses, and for WT versus ΔSO 10 min infection TMT experiments, vsn-normalization was performed grouped by sample type (phage genotype + timepoint) for phosphoproteomics; or, for proteomics, grouped by the protein’s species origin (host proteins normalized together over all times, phage proteins normalized by timepoint). Limma was used to compare samples to the uninfected timepoints, to compare genotypes at matched timepoints, and/or to compare timepoints across a specific phage infection. The model design included the phage genotype and timepoint status for phosphoproteomics, or considering sample type (phage genotype + timepoint) and replicate status for proteomics. The linear model was fitted using limma’s lmFit() function and a moderated t-statistic was computed using limma’s eBayes() function. P values were adjusted for multiple testing using the Benjamini–Hochberg procedure implemented in limma. Proteins were considered to be statistically significantly changed if they had adjusted P < 0.05 and absolute log2[FC] > 1.5. Phosphopeptides were considered to be statistically changed if they had adjusted P < 0.05 and an absolute log2[FC] > 1.
For the Retron-Eco9 and DarTG1 construct expression experiments, vsn-normalization was performed without grouping of samples/proteins. For the T7K SO domain solubility analysis, no vsn-normalization was performed. In both experiments, statistical significance was assessed using unpaired t-tests, using the stat_compare_means() function of the ggpubr R package.
For analysis of protein expression in BASEL phages from DIA data, protein groups with at least one unique peptide were considered, with only successful quantifications (no zero or NA values) considered further for normalization and plotting. vsn-normalization was performed without grouping of samples/proteins.
For the analysis of the TMT-multiplexed WT versus ΔSO kinetics (2, 3 and 4 min after infection) experiment, no normalization was performed. TMT reporter ion intensities were summed at the phosphopeptide (modified peptide) level for peptides with quantification in all three replicates of a given phage infection.
MS analysis and statistics (stoichiometry)
For the stoichiometry analysis, only peptides that spanned residues previously found to be phosphorylated in the time-course phosphoproteomics experiment were considered. The following PSMs were not considered for downstream analysis: (1) PSMs without complete TMT labelling (that is, without peptide N-terminal labelling); (2) PSMs modified with Ser/Thr/Tyr phosphorylation or M oxidation variable modifications; (3) PSMs mapping to proteins identified as differentially abundant in the time-course protein expression analysis (adjusted P < 0.05, FC of <0.5 or >1.5 at T = 10 min, WT T7 versus Δ0.7).
Two limma analyses were performed, and both analysed summed peptide intensities that were vsn-normalized without any sample/protein grouping. The first was performed to calculate and assess the quality of the phosphatase/control ratios for each phage genotype (WT T7, Δ0.7) and each of three biological replicates by inputting summed vsn-normalized TMT reporter intensities and comparing phosphatase versus control samples. It considered the sample (WT–phosphatase, WT–control, Δ0.7–phosphatase, Δ0.7–control) and replicate status in the model design. Differential analysis with limma was performed as described above. We considered some peptides to have poor phosphatase/control ratios when they were found in this analysis to be either significantly downregulated (limma, adjusted P < 0.05, log2[FC] < 0) or the log2[FC] was low (log2[FC] < −0.2, no significance requirement). These were excluded from the subsequent analysis. The estimated stoichiometry value for each genotype was calculated from these ratios using the following calculation: \(100\times (1-1/{2}^{{\log }_{2}[{\rm{FC}}]})\). The T7K-mediated stoichiometry was calculated by subtracting the estimated stoichiometry of the Δ0.7 from the WT T7 sample.
The second limma analysis was performed to determine differential stoichiometry, that is, stoichiometry deposited by the T7K. It was assessed by inputting manually calculated phosphatase/control ratios for each genotype and replicate for the peptides that had quality phosphatase/control ratios (see above) in both samples (T = 10 WT and T = 10 Δ0.7). These were manually calculated as: 2summed vsn-normalized phosphatase intensity/2summed vsn-normalized control intensity. The limma model was constructed to consider the phage genotype (WT T7 or Δ0.7) and replicate status. Peptides were considered to be of high differential stoichiometry if they had adjusted P < 0.05 and log2[FC] > 1.5.
GO analysis
To categorize phosphoproteins, goslim_prokaryotic_ribbon annotations were downloaded from QuickGO95 and annotated onto proteins. Proteins were assigned one of four representative GOSlim terms using the following ordered scheme: (1) transcription, translation (GOSlim terms: ribosome biogenesis, translation, regulation of DNA-templated transcription, DNA-templated transcription); (2) metabolic processes (GOSlim terms: metabolic process, primary metabolic process, lipid metabolic process, generation of precursor metabolites and energy, amino acid metabolic process, DNA metabolic process, small molecule biosynthetic process); (3) other (GOSlim terms: cell wall organization or biogenesis, detoxification, protein folding, transport, response to stimulus); or (4) unannotated, if it could not be mapped to a GOSlim.
GO-term analysis was performed using the STRINGdb R package (v.2.14.3)96, initialized with the following settings: species = 511145, version = 12.0, score threshold = 400. The get_enrichment() command (with defaults) was applied. For the non-phosphorylated proteome subset, host proteins (≥1 unique peptide identified in the parallel proteomics measurement) that were not found to have any uniquely mapping T7K-dependent phosphopeptides were compared to all detected host proteins (≥1 uniquely mapped peptide identified in the parallel proteomics measurement) as background. For the stoichiometry analysis, host proteins mapping to high differential stoichiometry peptides were considered, comparing to all detected host proteins (≥1 uniquely mapped peptide identified in the same experiment) as background.
Analyses with published phosphoproteomes
Phosphosites for E. coli (K-12) were downloaded from the dbPSP2.0 (ref. 97) in June 2024. Mass spectra from several previously published phosphoproteomics studies of E. coli (PRIDE submissions PXD008369, PXD008921 and PXD008289) were reanalysed as described above, with IDs harmonized using the same FASTA database.
The fraction of phosphorylated Ser, Thr and Tyr residues in different datasets was calculated per phosphorylated protein as the number of phosphorylated Ser, Thr and Tyr residues divided by the total number of Ser, Thr and Tyr residues present in the protein’s sequence. For the T7 dataset, either (1) all detected sites from the TMT time-course; or (2) only sites sensitive to T7K were considered. The phosphorylation fraction was also calculated for the human proteome based on (1) all phosphosites reported in the PhosphoSitePlus database98; (2) a comprehensive collection human phosphosites measured by MS curated in a recent stringent meta-analysis69; and (3) a single experimental dataset (using the same phosphoenrichment protocol as performed in this study) comprising phosphosites that were sensitive to EGF stimulation in HEK293F cells70. All distributions were compared to the distribution of T7-sensitive set of proteins using unpaired two-sample Wilcoxon signed-rank tests, implemented in the stat_compare_means() function of the ggpubr R package.
T7K homology search
To find putative homologues to the T7 kinase, we queried both viral and bacterial sequence databases using HMMs built from PHROG 2828 (ref. 99) (encompassing UniProt entry P00513 of T7K) augmented with the JSS1 kinase sequence. To distinguish the T7K kinase-domain and the shutoff-domain, we subset the PHROG 2828 alignment (generated using mafft100, v.7.505) into the parts corresponding to both domains (based on known domain architecture in T7 kinase), built an HMM for each and searched with both HMMs individually. To conduct the search, we used pyhmmer101 (v.0.10.15) and searched the databases against representative sequences of BFVD61 as well as against genomes from BASEL phage collection59 (to which we added the JSS1 genome) and all genomes from progenomes3 (ref. 62). All HMM hits were filtered using an E value of 10−10. We built a phylogeny (using fasttree102, v.2.2) on a multiple-sequence alignment of all identified T7K kinase domains. We did not consider hits coming from metagenome-assembled genomes as well as low-quality assemblies containing >400 contigs. To account for likely viral contamination, we reassigned contigs originally coming from progenomes3 to be of viral origin if they were longer than 40 kb and contained no bacterial marker gene (bacterial marker genes were identified using gtdb-tk103 (v.2.1.1). Out of 32 total contigs with putative T7K homologues found in progenomes3, we re-assigned 27 contigs as likely viral contaminants in this way. The remaining 5 contigs are all exceptionally long and contain multiple bacterial marker genes (Extended Data Fig. 9c). For those contigs, we ran PHASTEST104 (through the web API) to pinpoint probable prophage regions. We ran InterProScan (through the online interface) to identify protein domains on all hits. AlphaFold3 (ref. 105) predictions for several clade representatives with non-canonical domain configurations were also generated using the AlphaFold webserver.
Sequence motif analysis
For the analysis of T7K kinase substrate activity motifs, the ten amino acid residues surrounding each side of a phosphorylated residue (only from sites with PTMProphet localization score ≥ 0.75) were extracted and used to generate sequence logo plots using the ggseqlogo R package (v.0.2)106.
For the analysis of kinase active site motifs sequence variability, we extracted residues corresponding to the active sites of both viral and human kinases and subsequently generated sequence logo plots using the ggseqlogo R package (v.0.2)106 of the active sites and surrounding residues. Human kinase domain sequences for DFG motif generation were obtained by downloading the human kinome domains curated in the KinBase resource (http://kinase.com/web/current/kinbase/).
Protein structure analysis
Phage T7 protein structures were predicted by subjecting the Swissprot phage T7 proteome (UP000000840) sequences to AlphaFold3105 modelling through the AlphaFold Server. The PyMOL Molecular Graphics System (v.3.03 Schrödinger) was used to calculate the relative per-residue solvent accessible surface area on the top ranked model of each protein using the get_sasa_relative command. ChimeraX (v.1.10)107 was used to globally align (matchmaker) the AlphaFold3 prediction of T7K + Mg2+ + ATP with the Protein Data Bank (PDB) experimental structures 1FIN108 and 1KV2 (ref. 109).
FoldSeek110 structural homology searches were performed through the FoldSeek webserver, using the top-ranked AlphaFold3 (ref. 105) models generated for the full-length T7K and for T7K kinase-domain only (residues 1–242), and searching all available databases in 3Di/AA mode.
For the analysis of the T7K AlphaFold3 protein structure, PyMOL was first used to convert from .cif to .pdb format. This .pdb file was submitted to the GraphBind server111, with all ligand-specific options (including DNA and RNA) considered. Residues reported to be DNA or RNA binding are those that passed GraphBind’s internal binary thresholds.
Plasmid-based infection assays
To test a wide range of phage MOIs, serial dilutions of phage lysates (10 μl final volume) were added in round-bottom 96-well plates. We based MOI calculations on the assumption that an optical density of 1 corresponds to 1 × 109 bacterial cells. Plates containing the phages were pre-warmed at 37 °C for 1 h before adding the cells. E. coli BW25113 strains carrying plasmids encoding defence systems (with native promoters or with the inducible pBAD promoter; Supplementary Table 5), or a control plasmid, were diluted 1:2,000 from overnight cultures and grown at 37 °C. Strains with a plasmid with a native promoter were grown until OD600 = 0.05 (1 cm pathlength); strains with a plasmid with an inducible pBAD promoter were grown until OD600 = 0.6 (1 cm pathlength) in arabinose-containing medium (0.2% l-arabinose) as the pBAD promoter is catabolite repressed in LB during exponential growth. Note that, for Retron-Eco9 (Fig. 3a and Extended Data Fig. 5e), infection at OD600 = 0.05 (1 cm pathlength) led to inconsistent results, possibly because the culture entered stationary phase before the phage managed to collapse the bacterial population. We therefore started to dilute the cells 1:10 when they reached OD600 = 0.05, infecting the cells at OD600 = 0.005, which resulted in consistent phenotypes and effect sizes. Then, 90 μl of culture was added to the 96-well plates that contained phages and the plates were sealed with breathable membranes (Breathe-Easy). We kept several LB-filled wells in each experiment for background subtraction and to check for potential contamination. The plates were then incubated with constant shaking at 37 °C in a microplate reader, and the OD600 was measured every 10 min to acquire growth curves (either in a Biotek Synergy HT or Biotek synergy H1). Note that the OD600 values presented throughout all figures are from a microplate reader and were not pathlength corrected.
Infection assays with natural isolates
The natural isolate strain collection (536 unique strains, 558 strains in total; Supplementary Table 7) is organized on 6 different 96-well plates. We reorganized some strains so that each plate contained at least two empty wells to check for cross-contamination and to use for background subtraction. The whole collection was screened in four batches.
Phage lysate stocks were prepared to account for final MOIs of 5, 0.01, 0.001, 0.0001 and 0, determined for the reference strain BW25113, and assuming that an OD600 of 1 corresponds to 109 bacterial cells. We diluted the overnight cultures 1:2,000 and grew the strains at 37 °C and 800 rpm in deep 96-well plates until the first strains reached OD600 = 0.1 (1 cm pathlength). Note that the strains grow at different rates, which makes it impossible to infect them all at the exact same OD600 = 0.1 (1 cm pathlength). Thus, the exact phage to bacteria ratio will vary for each strain, but we tried to account for that by applying a range of MOIs. We then pipetted 90 µl of the bacterial cultures into 9 different round-bottomed 96-well plates, and added 10 µl of phage lysates on top (a single MOI per plate). Plates were sealed with breathable membranes (Breathe-Easy) and incubated without lids in a humidity-saturated incubator (Cytomat 2, Thermo Fisher Scientific) with continuous shaking. Plates were measured every 45 min in a Filtermax F5 multimode plate reader (Molecular Devices).
Out of the 558 strains, the uninfected control (MOI = 0) did not grow for six strains (ECOR-70, MG1655, ATCC 8739, S17 plate1, DH1, DH1 (E381)), so they were removed from the analysis. Strains with sequences that did not match the expected genotype were also removed from the analysis (all strains with a/b/c/x as a suffix in their name; Supplementary Table 7); note that 17 strains are present in replicates (same strain names) within the collection and we added the plate number behind the name to distinguish these in the analysis (Supplementary Fig. 3). One strain (NR-2654) appeared twice on the same plate, so we added _1 and _2 to the name, respectively. Three strains (IAI01, S17 and NR-2654) had replicates with different strain IDs (NT numbers), and we counted them only once among the unique strains. All replicates fell in the same categories (uninfected, infected without evidence of defence, infected with evidence of defence), except for three (MP1, ECOR-09, ECOR-10). For these strains, we removed the replicate from the uninfected category, as MP1 was validated from an independent glycerol stock to be infected (Supplementary Fig. 3), and ECOR-09 and ECOR-10 were slowly growing in the replicate that showed infection (Supplementary Fig. 3), so the other replicate was probably contaminated. For the analysis, we therefore assessed 513 unique strains (529 strains in total). Four other strains (BL21(DE3), DH5α, ECOR-09 plate 5, ECOR-10 plate 5) grew very slowly and the integration time for the AUC calculation (see below) was therefore increased from 6 to 10 h (Supplementary Fig. 3). The whole collection was screened in one replicate, and all 22 strains showing infection and evidence of defence (Supplementary Fig. 3c) were repeated from an independent glycerol stock and with higher MOI resolution (validations) in at least two biological replicates, that is, starting from two independent colonies on a plate (Supplementary Fig. 4). For the phage infection experiments comparing the different T7K mutants (Δ0.7, ΔSO, G76F) in natural isolate strains, we used an expanded MOI range (5, 1, 0.5, 10−1, 10−2, 10−3, 10−4, 10−5, 10−6, 10−7, 10−8 and 0).
Infection assay data processing
All data from plate reader assays were processed and plotted using MATLAB R2017b (The MathWorks). The data were background-subtracted with one background value (minimum measured value across each plate and timepoints) and plotted over time. For the natural-isolate screen, we subtracted the background at each timepoint with the minimum measured value across all plates of a batch at the timepoint. The AUC was determined under the linear growth curve from 0 h to 6 h of measurement using the function trapz. AUCs were normalized to the AUC of the uninfected control (MOI = 0) of the same strain (normalized AUC). To estimate the effect size for a biological replicate of a strain, we interpolated the MOI at an AUC of 0.5 (0.75 for ECOR-63 in Supplementary Fig. 5, as curves did not readily reach 0.5) on the log10 scale for both T7 WT and mutants, using the MATLAB function interp1. We then took the inverse logarithm base 10 of the difference (Δ0.7 minus WT). As the effect size strongly depends on well-matched phage titres of T7 WT and Δ0.7, we considered only a mean effect size of >5-fold or <0.2-fold as a hit (Fig. 4b and Extended Data Fig. 7a).
Bacterial isolate genome analyses
Genomic nucleotide sequences for the majority of the E. coli natural isolates were downloaded from the ECOREF53 resource’s website (https://evocellnet.github.io/ecoref/). The genomes of several strains (IAI09, IAI41, Perfeito-Clone-6, H10407, HM-341) were additionally sequenced using long-read sequencing. Overnight cultures of strains grown at 37 °C with shaking, were used for DNA extraction with the DNA mini-prep kit (ZymoBIOMICS). For library preparation, 1.2 μg of DNA for each sample was taken into fragmentation using Megaruptor 3.0 (Diagenode) with a speed of 31 and a final volume of 100 μl. Then, 1× clean-up with SMRTBell beads (PacBio) was performed. All recovered material was used as an input for library preparation using SMRTBell 3.0 chemistry (PacBio) and following the manufacturer’s instructions. Final libraries were quantified by Qubit and Femto Pulse system was used to assess the fragment distribution. Libraries were then pooled equimolarly and loaded in the Sequel IIe instrument at 100 pM with 30 h video time. The resulting PacBio reads were assembled into high-quality, complete genomes using Flye112 (v.2.9-b1768) using standard parameters.
To annotate known bacterial anti-phage defence systems in natural isolate strain genomes, nucleotide sequences were either (1) inputted directly into the PADLOC Python program (v.2.0.0, using the default settings)113 using PADLOC-DB v.2.0.0; or (2) first translated into a protein FASTA file using the Prodigal Python package (v.2.6.3, using the default settings)114, and then inputted into the DefenseFinder python package (v.1.2.2, using the default settings)6 using DefenseFinder models (v.1.2.3)115 and CasFinder models (v.3.1.0)116.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.