Selection and prioritization of extracellular tumour-derived genes
Candidate extracellular tumour-derived genes were identified using a stepwise filtering and prioritization strategy integrating single-cell transcriptomic analyses, protein subcellular localization annotations, functional dependency data, domain-based functional annotation and literature-based curation (Extended Data Fig. 1a).
Publicly available scRNA-seq datasets from human14 and mouse15 PDAC were independently analysed. Within each dataset, malignant or premalignant epithelial populations were defined. Differential gene expression analysis was performed comparing malignant or premalignant cells to non-malignant epithelial cells. Genes were retained if they met stringent significance and effect size thresholds (adjusted P value ≤ 0.05, log2 fold change > 4 for mouse datasets and adjusted P value ≤ 0.05, log2 fold change > 3 for human) and showed predominant expression within malignant or premalignant epithelial compartments (Extended Data Fig. 1a). Gene lists derived from human and mouse datasets were subsequently merged to generate an initial cross-species candidate pool.
To enrich for factors with extracellular activity, protein subcellular compartment annotations were obtained from the Compartments database53. Enrichment analysis was performed, and genes annotated to the enriched term ‘extracellular region part’ were retained. These genes were intersected with the differentially expressed gene sets to further refine the candidate list.
To exclude genes whose perturbation would be expected to strongly impair tumour cell viability under standard in vitro conditions, candidates were evaluated using CRISPR–Cas9 dependency scores from the DepMap database across 46 human pancreatic cancer cell lines. Mean CRISPR dependency scores were calculated for each gene, and genes with average scores between −0.15 and 0.15 were retained, indicating minimal effects on cell-intrinsic fitness (Extended Data Fig. 1b).
Pathway- and domain-level analyses were performed to characterize the biological programmes represented within the candidate pool (Extended Data Fig. 1c,d). These enrichment analyses were used to provide biological context and to guide a literature-informed selection of the final gene set (Extended Data Fig. 1e).
The selection rationale for each gene is summarized in Supplementary Table 1.
Lentiviral vector construction and production
A comprehensive protocol for generating PC/CRISPR lentiviral vectors and library pools is available on the Addgene website, as part of the Pro-Code vector kit (Addgene 1000000197) and has been described previously12,54. Below, we provide a brief overview of the procedure. For selecting sgRNA sequences targeting each gene, we used the Brie CRISPR library for mouse genes55. A complete list of the oligonucleotides used for the experiments in this paper is provided in Supplementary Table 3. Library cloning was performed using nuclear PC/CRISPR lentiviral vectors (Addgene 1000000197), while validations utilized lentiCRISPR v2 (Addgene 5296156). Library cloning was conducted in a 96-well plate format, where constructs were annealed, ligated, and transformed individually but processed in parallel, as outlined in the Pro-Code vector kit protocol. Oligonucleotides were prepared by resuspending them to 100 µM in water. For annealing, forward and reverse oligos were mixed to a final concentration of 2 µM, combined with 10× NEBuffer 2.1 and water. Next, they were heated to 95 °C for 5 min, followed by gradual cooling to room temperature. The PC/CRISPR lentiviral vector was digested with BbsI-HF (NEB), per the manufacturer’s instructions, and purified using Qiagen PCR purification columns. Annealed oligos were ligated into the digested vector backbone by combining 50 ng of digested plasmid with 6–8 ng of annealed oligos, incubating at room temperature for 10 min with Quick Ligase (NEB). Next, 5 µl of the ligation reaction was transformed into 50 µl of TOP10 chemically competent bacteria. After incubating on ice for 30 min, the bacteria were heat-shocked at 42 °C for 30 s, cooled on ice for 2 min, and plated on LB ampicillin agar plates overnight at 37 °C. Colonies were picked, cultured, and plasmid DNA was extracted using the Zymo ZR Plasmid MiniPrep Classic Kit. The sgRNA sequence was confirmed via Sanger sequencing. For cloning into the lentiCRISPR v2 backbone, a similar process was followed. However, sgRNAs were inserted into the BsmBI site.
Lentiviral vector production was performed as previously described57 and detailed in the Pro-Code_Kit_Methodology.pdf on the Addgene website. In brief, HEK293T cells were seeded at 500,000 cells per well in 6-well plates and incubated at 37 °C with 5% CO2. After 24 h, cells were transfected using calcium phosphate with third-generation lentiviral packaging plasmids and the transfer plasmid (pVSV (1 µg), pMDLg/pRRE (2 µg), pRSV-REV (1 µg), and PC/CRISPR vector (6 µg)). The plasmids were first mixed with 2.5 M CaCl2, vortexed, and incubated for 10 min. Then, 2× HBS solution (281 mM NaCl, 100 mM HEPES, 1.5 mM Na2HPO4, pH 7.05) was added dropwise with gentle vortexing. The transfection mixture was applied to cells, and the medium was replaced after 14 h. Supernatants were collected 30 h after medium replacement, filtered through a 0.22 µm PVDF disc filter, and stored at −80 °C.
Cell culture
FC1245 PDAC cells (KPC cells, described above) were generated from a primary tumour in a KrasLSL-G12D/+;Trp53LSL-R172H/+;Pdx1-cre mouse and were provided by D. Tuveson. 8442 PDAC cells (KC cells, in this manuscript) were generated from a primary tumour in a KrasLSL-G12D/+;Ptf1acre/+ mouse and were provided by D. Saur.
These cells, and all lines resulting from their genetic modification, were routinely passaged in DMEM (Thermo Fisher Scientific) containing 10% heat-inactivated FBS and 100 U ml−1 penicillin/streptomycin, for no more than 25 to 30 passages. Pro-Code+ cells were verified by flow cytometry (for mCherry positivity, as readout for Pro-Code presence in the culture) and mass cytometry (for their sgRNA/Pro-Code identities) before every in vivo experiment. The murine PDAC cell lines were authenticated by genotyping PCR. All cell lines were routinely tested for mycoplasma contamination and tested negative.
To determine doubling time of cell lines, 1,000 cells per well were seeded in 100 μl of growth medium in technical triplicates in 96-well plates. After 24 h intervals, cells were fixed and stained with 0.2% Crystal Violet in an ethanol:water solution. Crystal Violet was solubilized with 10% acetic acid and absorbance was quantified at 595 nm. The resulting values were used to determine doubling times. Experiments were performed in three biological replicates.
Ex vivo BMDM culture
BMDMs were generated from bone marrow isolated from C57BL/6J mice using established protocols58. In brief, bone marrow was flushed using cold sterile PBS and RBC-lysed for 1 min at room temperature. Cells were plated in DMEM containing 10% v/v FBS and 10 ng ml−1 recombinant macrophage-colony stimulating factor (M-CSF; PeproTech, 315-02). Cells were plated at a concentration of ~150,000 cells per cm2 on non-treated Petri plates. At day 2, medium was replenished 1:1 with fresh medium containing 10 ng ml−1 M-CSF. At day 6, cells were gently detached using ice-cold PBS containing 5 mM EDTA and replated onto test plates.
Fibrin co-culture assays
KPC cells were plated in 10-cm dishes and cultured for 4 days, after which conditioned medium was collected, filtered through 0.2-µm pore filters, and used for downstream assays. 12-well plates were coated with fibrin by combining fibrinogen and thrombin solutions (Millipore Sigma, ECM630) according to the manufacturer’s instructions. BMDMs (3 × 105 per well) were plated either onto fibrin-coated wells or onto uncoated plastic in the presence of KPC conditioned medium. Where indicated, cultures were treated with the thrombin inhibitor Argatroban monohydrate (Selleckchem, S5074) at a concentration of 10 µM. After 24 h of incubation, cells were detached and prepared for flow cytometry.
Vector transduction
For the Perturb-map experiment, KPC cells were transduced with PC/CRISPR lentiviral vectors. Cells were seeded in 12-well plates at a density of 20,000 cells per well, 24 h prior to transduction. The next day, lentiviral vectors were added to the cells in the presence of 5 µg ml−1 polybrene (Millipore) at a low multiplicity of infection (MOI), with each well receiving a distinct PC/CRISPR vector. The transduced cell populations were pooled based on their mCherry+ percentage to achieve an equal distribution of PC/CRISPR populations. Cells expressing PC/CRISPR were sorted for mCherry positivity, ensuring >99% purity. Subsequently, KPC cells were transduced with Cas9 lentivirus and selected using 4 µg ml−1 puromycin. PC/CRISPR cells, both with and without Cas9, were maintained in culture with and without puromycin for two weeks before proceeding with downstream analyses.
A similar procedure was used to transduce KPC cells with the lentiCRISPR v2 lentiviral vectors used for validation experiments.
Mouse experiments
All animal studies were conducted in accordance with the ARRIVE guidelines and were approved by the Institutional Animal Care and Use Committee (IACUC) of the Icahn School of Medicine at Mount Sinai under the applicable institutional animal protocol. Both male and female mice were used. Cell lines derived from female autochthonous tumours were transplanted into female recipients, whereas cell lines derived from male autochthonous tumours were transplanted into male recipients.
Cas9 expressing mice (strain 028239), Rag2−/− mice (strain 008449), C57BL/6J mice (strain 000664) or NSG mice (strain 005557) obtained from the Jackson Laboratory, were used at 8–10 weeks of age. All animals were kept in a dedicated facility, under a 12 h light:12 h dark cycle, a housing temperature between 20 and 24 °C and a relative air humidity of 55%.
No statistical methods were used to predetermine sample sizes. Sample sizes were selected based on prior experience with the experimental systems, the expected variability of the orthotopic tumour models and the number of biological replicates required to assess reproducibility while limiting animal use. Individual mice were considered biological replicates, and exact sample sizes are reported in the corresponding figure legends. No animals or data points were excluded from the analyses.
To generate orthotopic pancreatic tumours, 50,000 KPC or KC cancer cells were grafted into the pancreas of 8–10-week-old mice following established protocols47. Transplantation efficiency of all lines used was consistently high and comparable across groups, and we did not observe significant differences in engraftment across conditions.
For animal treatment studies initiated after tumour implantation, mice were randomly allocated to treatment groups. For macrophage-depletion and CD18-blockade experiments, which required treatment before tumour implantation, mice were prospectively allocated to the indicated treatment and matched control groups before treatment initiation. Random allocation was not applicable to experiments comparing genetically distinct tumour cell populations because group assignment was determined by the transplanted cell line or genetic perturbation. In pooled Perturb-map experiments, all mice received the same pooled cell population, and allocation to separate experimental groups was therefore not required.
Investigators administering treatments were not blinded to group allocation because the treatment regimens, routes of administration and schedules differed between groups. Where feasible, sample collection and downstream analyses were performed blinded to experimental group, including tumour mass measurements, immunohistochemical quantification and flow cytometry analysis.
For ICB, ten days after cancer cell transplantation, mice were randomized into two groups. Randomized mice were injected intraperitoneally with 200 μg per dose of InVivoMAb anti-mouse PD-1 (CD279) Clone 29F.1A12 (BioXcell, BE0273) or InVivoMAb rat IgG2a isotype control (BioXcell, BE0089) antibodies, every third day.
For PAI1 inhibition studies, ten days after tumour cell injection, mice were randomized into four treatment groups. The first group received PAI-039 (Selleckchem, S7922; 20 mg kg−1, 5 days a week) by oral gavage; the second group received intraperitoneal injections of InVivoMAb anti-mouse PD-1 (CD279; clone 29 F.1A12, BioXcell, BE0273) at 200 µg per dose every third day; the third group received a combination of PAI-039 and anti-PD-1; and the fourth group received InVivoMAb rat IgG2a isotype control antibodies (BioXcell, BE0089) together with the PAI-039 vehicle. PAI-039 was prepared in a vehicle consisting of 5% DMSO and 95% corn oil.
For macrophage depletion experiments, mice were divided into two groups: control and macrophage depletion. Control group mice received a single dose of 1 mg IgG (InVivoMAb rat IgG1 isotype control, anti-trinitrophenol, BioXcell, BE0290) on day 1, followed by 200 µl of control liposomes (PBS) (Encapsula, CLD-8914) on day 2. In the macrophage depletion group, mice were treated with 1 mg of InVivoMAb anti-mouse CSF1 (BioXcell, BE0204) on day 1 and 200 µl of clodronate liposomes (Encapsula, CLD-8914) on day 2. After one week, mice underwent orthotopic transplantation with KPC control, Serpinb2-KO or Serpine1-KO cells. Next, four additional treatment cycles were administered every third day, with each cycle consisting of 0.5 mg of IgG or anti-CSF1 on the first day, followed by 200 µl of control or clodronate liposomes on the second day for the control and macrophage depletion groups, respectively. The experiment concluded after two weeks from tumour cell injection.
For CD18-blocking experiments, mice were divided into treatment groups receiving isotype control, CD18 blockade alone, or combined CD18 and PD-1 blockade. Mice assigned to CD18 blockade were treated with InVivoMAb anti-mouse CD18 antibody (Rat IgG2a, BioXcell, BE0009) at a dose of 10 mg kg−1 via retro-orbital injection (100 µl per injection) twice per week. Treatment was initiated one week prior to orthotopic tumour cell transplantation and continued for two weeks following transplantation until the experimental end-point. For combination treatment experiments, mice additionally received 200 µg of InVivoMAb anti-mouse PD-1 (CD279) antibody (clone 29 F.1A12, BioXcell, BE0273) by intraperitoneal injection every third day following tumour transplantation. Control mice received matched injections of InVivoMAb rat IgG2a isotype control, anti-trinitrophenol (BioXcell, BE0089) according to the corresponding treatment schedules and routes of administration.
Animals were euthanized at the predefined experimental time points specified in the figure legends or upon reaching a humane end-point in survival experiments. Humane end-points included loss of more than 20% of initial body weight, compromised mobility or other clinical signs of distress. No animal exceeded the approved humane end-points.
Flow cytometry
For cell sorting experiments, adherent cells were detached with 0.05% trypsin-EDTA, washed in PBS and resuspended in cell culture medium. Samples were sorted on the BD FACSAria III Sorter (BD Biosciences).
For flow cytometry profiling of the TME, fresh PDAC samples were minced and enzymatically digested with the tumour dissociation kit (Miltenyi, 130-096-730) for 40 min at 37 °C with agitation. The cell suspension was strained through a 70 µm strainer, spun down and resuspended in flow cytometry buffer (PBS, 2% bovine serum albumin, 5 mM EDTA). Cells were centrifuged at 350g for 5 min at 4 °C and then pellets were resuspended with ACK lysis buffer (Life Technologies) to lyse red blood cells at room temperature for 10 min and washed with cold flow buffer. Samples were then resuspended in flow cytometry buffer and stained for 30 min at 4 °C. All antibodies used are detailed in Supplementary Table 4. Upon staining, cells were analysed using a BD LSR Fortessa. Flow cytometry data were acquired using the FACS Diva software v.7 (BD) and were analysed using FlowJo (v10.9.0). Absolute cell numbers were calculated using the initial sample volume and fluorescent counting beads (AccuCheck Counting Beads PCB100, Molecular Probes), according to the manufacturer’s instructions, and normalized to the initial tumour mass to obtain cell numbers per milligram of tumour.
CyTOF mass cytometry
Cell suspension processing and CyTOF analysis were carried out as described previously54. In brief, 3 × 106 cells were collected, resuspended in PBS and stained for viability using Cell-ID Intercalator-103 Rh for 15 min at 37 °C. Surface marker staining was then performed in flow buffer with an anti-mouse CD16/CD32 blocking antibody (eBioscience) on ice for 30 min. Cells were subsequently fixed and permeabilized using the eBioscience FOXP3/Transcription Factor Staining Buffer Set (Invitrogen) following the manufacturer’s instructions. Afterward, cells were stained with epitope-tag antibodies on ice for 1 h and incubated with 125 nM Ir intercalator (Fluidigm) diluted in PBS with 2.4% formaldehyde at room temperature for 30 min. Following staining, cells were washed and stored in 10% DMSO FBS at −80 °C until acquisition. Samples were acquired using either a CyTOF2 or Helios instrument (both from Fluidigm) at an event rate of <500 events per second. Antibodies were purchased in purified form and conjugated in-house using MaxPar X8 Polymer Kits (Fluidigm) according to the manufacturer’s protocol. Details of the antibodies used for mass cytometry can be found in Supplementary Table 4.
CyTOF data analysis
CyTOF data analysis was conducted as previously detailed54. In summary, manual gating was performed on Cytobank (web platform) to select single, live, and Pro-Code-positive (mCherry+) cells. Pro-Code-positive cells were subsequently debarcoded using the Single Cell Debarcoder tool59.
Western blot
After cell culture, the medium was removed, and cell culture plates were washed twice with ice-cold PBS. Cells were lysed using 150 μl of RIPA buffer (Thermo Scientific, 89900) containing protease and phosphatase inhibitors. Lysis and lysate collection were carried out on ice. The lysates were incubated on ice for 5 to 10 min before being transferred to 1.5 ml Eppendorf tubes and centrifuged at 18,000g for 10 min at 4 °C to collect the supernatant. Protein concentrations were quantified using the Pierce BCA Protein Assay Kit (Thermo Scientific, 23227) per the manufacturer’s protocol. For electrophoresis, 30 µg of protein from each sample was loaded per lane along with the PageRuler Plus Prestained Protein Ladder (Thermo Scientific, 26619). Proteins were separated on an Invitrogen NuPAGE 10% Bis-Tris gel (Thermo Scientific, NP0315BOX) at 100 V. Subsequently, proteins were transferred onto methanol-activated PVDF membranes at 300 mA for 2 h. The membranes were blocked in 5% milk dissolved in TBS-T, tris-buffered saline (Fisher Bioreagents, BP24711), with 0.1% Tween-20 (Thermo Scientific, BP337) for 1 h at room temperature on a plate rocker. After blocking, membranes were washed three times with TBS-T (10 min per wash) and incubated overnight at 4 °C with the indicated primary antibodies. The antibodies were diluted in the blocking buffer according to the manufacturer’s instructions. The following day, primary antibodies were removed, and membranes were washed three times with TBS-T (5 min per wash) on a plate rocker. Membranes were then incubated with secondary antibodies diluted 1:10,000 in TBS-T for 1 h at room temperature with gentle rocking. Following secondary antibody incubation, membranes were washed three more times with TBS-T (10 min per wash) and incubated with chemiluminescence reagent (Pierce ECL Western Blotting Substrate, Thermo Scientific, 32209) per the manufacturer’s instructions. The following primary antibodies were used at a 1:1,000 dilution: anti-PAI1 (Invitrogen, MA1-40224), anti-PAI2 (Invitrogen, PA5-27857), and anti-Vinculin (Sigma-Aldrich, V4505). Secondary antibodies included anti-rabbit-HRP (Cell Signaling, 7074S) and anti-mouse-HRP (Cell Signaling, 7076P2). Full scans are provided in the Supplementary Data.
MICSSS
Multiplexed immunohistochemical consecutive staining on single slide (MICSSS) was carried out following a previously described protocol60. Where indicated, this workflow was applied to formalin fixed, paraffin-embedded (FFPE) sections from orthotopic mouse PDAC tumours and to de-identified archival FFPE tissue sections from 30 human PDAC cases obtained from the Mount Sinai Department of Pathology tissue repository. Human samples were provided without patient identifiers or linked clinical information. In brief, 5-µm-thick FFPE tissue sections were baked at 60 °C overnight, deparaffinized using xylene, and gradually rehydrated through a series of ethanol solutions (100%, 90%, 70% and 50% in water). Antigen retrieval was achieved by incubating the slides in Antigen Retrieval Solution (pH 9, Dako) at 95 °C for 30 min. The slides were then cooled to room temperature for 30 min, rinsed with TBS, and treated with 3% hydrogen peroxide at room temperature for 15 min to inhibit endogenous peroxidase activity. Following this, slides were blocked with Serum-Free Protein Block (Dako) for 30 min at room temperature and incubated with primary antibodies diluted in Antibody Diluent, Background Reducing (Dako) for 1 h at room temperature. After washing with TBS containing 0.04% Tween-20, HRP-conjugated secondary antibodies were applied for 30 min based on the species of the primary antibody, including EnVision+ System–HRP Labelled Polymer Anti-mouse (Dako), EnVision+ System–HRP Labelled Polymer Anti-rabbit (Dako), VisUCyte HRP Polymer Goat IgG Antibody (R&D Systems) or ImmPRESS HRP Anti-Rat IgG, Mouse Absorbed (Vector Laboratories). Antigen detection was carried out using the AEC Peroxidase Substrate Kit (Vector Laboratories), and counterstaining was achieved with Harris Modified Hematoxylin Solution (Sigma-Aldrich). The slides were then mounted with Glycergel Mounting Media (Agilent) and scanned at 20× magnification using the Aperio AT2 slide scanner (Leica). For additional staining rounds, the coverslips were removed by immersing the slides in 60 °C water. Residual AEC and haematoxylin were stripped using a sequential treatment with ethanol solutions (50%, 70% (containing 1% HCl 12N), and 100%). Slides were then processed according to the original protocol with one modification: an additional blocking step was introduced. Depending on the species of the primary antibody used for that round, the slides were incubated for 30 min at room temperature with one of the following Fab fragment reagents: AffiniPure Fab Fragment Donkey Anti-Mouse IgG (H + L), AffiniPure Fab Fragment Donkey Anti-Rabbit IgG (H + L), AffiniPure Fab Fragment Donkey Anti-Goat IgG (H + L), or AffiniPure Fab Fragment Donkey Anti-Rat IgG (H + L) (Jackson ImmunoResearch). The antibodies utilized for MICSSS are detailed in Supplementary Table 4.
MICSSS image processing
The MICSSS data were processed and analysed using Fiji (v1.0) and QuPath (v0.5.1), following a previously specified protocol60. In brief, sequentially acquired images from each staining round were aligned using the Linear Stack Alignment tool with SIFT registration in Fiji. For each image, the haematoxylin and marker staining signals were separated through deconvolution in Fiji, utilizing the default Hematoxylin/AEC colour vector. Stacked layers of AEC signals for each staining round, along with one haematoxylin image, were pseudocoloured and combined to generate composite images. Cell segmentation was performed in QuPath using nuclear detection on the haematoxylin stain with optimized settings. The same segmentation parameters were applied consistently across all images within the experiment.
Pro-Code debarcoding on MICSSS data
Pro-Codes were assigned to cells using a modified algorithm based on Zunder et al.59, as previously described in detail12. In summary, the mean pixel intensity for each epitope tag within the nuclear regions of segmented cells was normalized to a scale of 0 to 1 for each tissue section. Epitope tags were then ranked by their normalized intensity for each cell, and the difference between the 3rd and 4th highest-intensity tags, referred to as the delta value, was calculated. Cells were assigned a Pro-Code corresponding to the three highest-intensity tags if their delta value exceeded 0.01. These Pro-Code assignments were cross-referenced with the vector library design to identify the associated sgRNA target genes. Pro-Code assignments were further validated by applying marker-specific intensity thresholds, which were predefined for each epitope based on optimization across tissue sections. Cells that did not meet the intensity thresholds for all three markers within a Pro-Code combination were excluded from the final assignment. All images were then combined in a large single-cell object and further analysed using Squidpy61 and visualized in R (v4.2.2).
Enrichment and depletion analysis
To assess enrichment or depletion relative to the internal control of the library, we first calculated the number of Pro-Code+ cells for each mouse. The Pro-Code+ cell counts were normalized against the total number of Pro-Code+ cells within each mouse to account for variations in overall cell numbers. Individual animals were treated as biological replicates to ensure statistically robust analysis. We applied multiple Mann–Whitney tests to compare each condition against the control, using the null hypothesis that no significant differences existed. P values were adjusted for multiple comparisons using the Benjamini–Krieger–Yekutieli FDR correction.
Tumour clonality assessment on MICSSS data
To assess tumour clonality on MICSSS data, we analysed spatially resolved cell coordinates of Pro-Code+ cells. We constructed k-nearest neighbour graphs to quantify clonal purity by calculating the fraction of neighbouring cells with mismatched barcodes. Dense clonal regions, or focal areas, were identified using mean-shift clustering with a 250 μm bandwidth, retaining clusters with at least 50 cells. Clonal diversity was quantified within these focal points using the Shannon diversity index (SDI) and the evenness index, capturing both heterogeneity and uniformity. Temporal and spatial variations in clonality were analysed by calculating median SDI and evenness across conditions, highlighting dynamic changes over time.
Neighbourhood enrichment analysis
To investigate spatial relationships between cell clusters across tumour tissues, we performed neighbourhood enrichment analysis using Squidpy (v1.3.0). The neighbourhood enrichment scores were computed using the squidpy.gr.nhood_enrichment function based on a spatial graph constructed via Delaunay triangulation, followed by radius- and percentile-based pruning. The enrichment statistic was computed using a permutation-based test to assess whether specific clusters occurred as neighbours more frequently than expected by chance. Higher scores indicate clusters that preferentially co-localize within the tissue. Conversely, clusters with low scores are spatially segregated, suggesting depletion. We used 10,000 permutations to ensure robust statistical assessment, unless otherwise specified. The enrichment scores (z-scores), for Pro-Code analysis and Perturb-map at days 7, 14 and 21, were extracted and saved for further interpretation and visualization.
MICSSS spatial distances
Spatial distance analysis between specific phenotypes was performed using the squidpy.tl.var_by_distance function in Squidpy (v1.3.0). For each cell of a particular phenotype within the defined ROIs, the shortest distance to a designated anchor point was calculated at the single-cell level. The resulting distances were visualized as distribution plots in the corresponding figure panels.
CyCIF
FFPE tissue sections were prepared and stained using cyclic immunofluorescence (CyCIF) as previously described62,63, with minor modifications. Tissue sections cut to 5 μm thickness on SuperFrost Plus II positively charged slides were baked overnight to promote tissue adhesion, followed by deparaffinization with three 5-min washes in xylene and rehydration through a graded ethanol series (100%, 90%, 70% and 50%; 5 min each) and then two washes in distilled water. Heat-induced epitope retrieval was then performed in a Tris-EDTA based 1× Target Retrieval Solution (DAKO, pH 9) at 95 °C for 30 min, followed by cooling at room temperature for 30 min and a brief wash in PBS.
To reduce tissue autofluorescence, sections were photobleached by immersion in bleaching solution (4.5% H2O2, 20 mM NaOH in PBS) under LED illumination for 2× 45 min and washed again in PBS. Slides were then incubated overnight at 4 °C with fluorophore-conjugated secondary antibodies (anti-mouse, anti-rat, and anti-rabbit; 1:1,000 dilution; see Supplementary Table 4) in SuperBlock (ThermoFisher) to assess non-specific binding and residual autofluorescence. After washing (3× 5 min PBS), slides were again photobleached by immersion in bleaching solution (4.5% H2O2, 20 mM NaOH in PBS) under LED illumination for 2× 45 min and washed again in PBS.
For each CyCIF cycle, slides were incubated with Hoechst 33342 (1:10,000; Thermo Fisher Scientific) and either fluorophore-conjugated primary antibodies or unconjugated primary antibodies diluted in SuperBlock buffer (see Supplementary Table 4). Primary antibody incubation was performed either overnight at 4 °C or for 2 h at room temperature in the dark. For unconjugated primaries, this was followed by a 2 h incubation with fluorophore-conjugated secondary antibodies at room temperature. Prior to image acquisition, slides were washed (3× 5 min PBS), mounted in 70% glycerol, and coverslipped with 24 × 50 mm no. 1.5 coverslips (Epredia).
Images were acquired on a RareCyte CyteFinder II HT automated microscope using the UV, AlexaFluor488, Sytox, and AlexaFluor647 detection channels with a 20× objective (NA 0.75) and 2 ×2 binning resulting in a final resolution of 0.65 µm per pixel. Exposure times were optimized for each channel to avoid saturation and kept constant across treatment groups. Following imaging, coverslips were removed by incubating slides in 1× PBS at 57 °C for 10 min. Between successive staining cycles, slides were photobleached (2× 45 min) as before and washed (3×5 min PBS). Cycles of staining, imaging, and bleaching were repeated until all antibody stains were acquired per panel.
CyCIF image pre-processing and quality control
Preanalytical CyCIF image processing, including stitching, image registration, illumination correction, segmentation and single-cell feature extraction, was performed using the MCMICRO pipeline (v1.0), an open-source modular microscopy workflow (https://github.com/labsyspharm/mcmicro). For generation of nuclear probability maps, a trained U-Net model (UnMicst v2) was applied, followed by marker-controlled watershed segmentation for single-cell identification. A diameter range of 3–5 pixels was used for nuclei detection. Probability maps generated by UnMicst were processed with S3segmenter to generate nuclear segmentation masks. Cytoplasmic regions were approximated by expanding the nuclear masks by three pixels. Mean fluorescence intensities for each marker were then calculated for every segmented cell, resulting in a single-cell feature table for each whole-slide CyCIF image. The xy coordinates of annotated histological regions were used to extract quantified single-cell data for cells located within the defined ROIs.
Multiple quality control steps were applied to ensure the accuracy of the single-cell data. At the image level, cross-cycle image registration and tissue integrity were visually inspected, and regions with poor registration, tissue deformation, or imaging artifacts were excluded from downstream analyses. Antibodies that produced low-confidence staining patterns upon visual inspection were excluded. Segmentation quality was evaluated iteratively, and segmentation parameters were adjusted to optimize the accuracy of the segmentation masks.
CyCIF single-cell phenotyping
Background intensity distributions for each marker were manually estimated on individual slides during manual gating using Gater (MCMICRO). The resulting marker-specific gates were used to rescale single-cell intensities between 0 and 1 using the rescale function implemented in scimap (v2.2.11), such that values > 0.5 indicated marker-positive cells. This procedure was applied independently to each image to account for slide-to-slide variability, after which all images were merged into a combined single-cell dataset. The scaled single-cell data were then used for cell-type annotation.
Cells were assigned to phenotype classes based on marker expression using a predefined Boolean phenotyping workflow. Each cell was classified according to the presence or absence of specific marker combinations defined in a marker relationship chart. Cells that did not match the Boolean criteria for any phenotype were assigned to an ‘unknown’ class.
CyCIF data analysis and Pro-Code calling
For each Pro-Code epitope channel, background intensity distributions were manually estimated on individual slides during manual gating using Gater (MCMICRO). To mitigate inter-slide variability in fluorescence intensity across Pro-Code epitopes, we applied a per-slide robust normalization strategy based on a Winsorized interquartile range (IQR). For each slide independently, epitope intensities were first thresholded to cap extreme negative values at a small constant to stabilize scaling. Slide-specific robust statistics were then calculated for each channel, including the median intensity and a Winsorized IQR defined as the difference between the 90th and 10th percentile values. This percentile range excluded extreme tails caused by rare hyper-bright artifacts while preserving sufficient dynamic range to distinguish negative and positive populations. Background correction performed upstream was retained during scaling.
Normalized epitope intensities were then used for delta-based triplet assignment to identify Pro-Code combinations. This normalization minimized fluorophore-specific brightness differences and inter-slide staining variability while preserving relative expression relationships between epitopes. Assigned Pro-Code identities and corresponding gene knockouts were stored in an AnnData object, and all downstream spatial analyses were performed using Squidpy.
ROI-based spatial mixing analysis
To quantify spatial mixing between Control (F8-KO, serpin wild type) and Serpine1-KO tumour populations, we implemented a region-of-interest (ROI)-based spatial analysis using single-cell coordinates obtained from CyCIF imaging. Tumour regions were first defined independently for each mouse by constructing alpha-shape polygons from the spatial coordinates of control-KO and Serpine1-KO-positive tumour cells. When multiple tumour lobes were detected, polygons smaller than 5% of the largest tumour area were excluded. To minimize edge artifacts, cells located within a 150-pixel border from the tumour boundary were excluded from candidate ROI centres.
Circular ROIs (radius = 200 pixels) were then generated by centring ROIs on candidate tumour cells. ROIs were retained if they contained at least 400 total cells and at least 200 tumour cells belonging to either the control or Serpine1-KO. Within each ROI, tumour composition was quantified as the fraction of the minority genotype among all control- and Serpine1-KO-positive tumour cells. ROIs were classified as low-mixing if the minority genotype represented ≤20% of tumour cells and high-mixing otherwise. The majority genotype (control or Serpine1-KO) was recorded for each ROI. In addition, the abundance of all annotated cell phenotypes within each ROI was quantified as both cell counts and fractions relative to the total number of cells.
To ensure spatially distributed sampling and reduce redundancy from overlapping ROIs, a farthest-point sampling strategy based on ROI centre coordinates was applied independently for each mouse. Up to 20 ROIs were selected for each of three categories: low-mixing ROIs with Control majority, low-mixing ROIs with Serpine1-KO majority, and high-mixing ROIs irrespective of majority genotype. The resulting ROI table was used for downstream analyses of TME composition.
Xenium spatial transcriptomics and post-Xenium CyCIF
To generate a multimodal transcriptomic-proteomic screening dataset, a custom Xenium gene panel was designed using the 10x Genomics Xenium Custom Panel Designer to target genes associated with tumour cell states, immune populations, stromal interactions, and interferon and inflammatory signalling pathways. Gene selection was informed by prior scRNA-seq data and published spatial transcriptomic datasets13,35,64,65,66,67. The final panel comprised 480 genes, including probes for mCherry–tdTomato detection of Pro-Code-transduced cells, markers of epithelial differentiation, immune activation, cytokine signalling, and stress-response programmes. All probes were synthesized and quality-controlled by 10x Genomics. A complete list of probes included in the Xenium panel is provided in Supplementary Table 5.
Xenium assay workflow
Spatial transcriptomic profiling was performed using the Xenium In Situ platform (10x Genomics) according to the manufacturer’s protocol for FFPE tissues, with minor modifications. In brief, tissues were sectioned onto a single Xenium slide within the custom fiducials (12 mm × 24 mm). Sections were deparaffinized, rehydrated, and subjected to target retrieval and protease digestion to optimize probe accessibility. Custom gene-specific probe sets were hybridized in situ, followed by rolling circle amplification and iterative fluorescent imaging cycles to detect individual RNA molecules at subcellular resolution.
Imaging was performed on the Xenium Analyzer using default acquisition settings. Cell segmentation was carried out using the Xenium onboard pipeline, leveraging nuclear staining and spatial transcript density to define cellular boundaries. Transcript molecules were assigned to segmented cells based on spatial overlap.
Post-Xenium cyclic immunofluorescence
Following transcriptomic imaging, slides were incubated in PBS until further processing. Within 48 h, residual 10x Genomics Autoquencher Reagent was removed using three 1-min washes with 10 mM sodium hydrosulfite in ddH2O to enhance imaging signal. After washing (3 ×5 min PBS), heat-induced epitope retrieval was performed in Tris-EDTA-based 1× Target Retrieval Solution (DAKO, pH 9) at 95 °C for 30 min, followed by cooling at room temperature for 30 min and washing in PBS.
To reduce tissue autofluorescence, sections were photobleached by immersion in bleaching solution (4.5% H2O2, 20 mM NaOH in PBS) under LED illumination for two 45-min intervals, followed by PBS washes. Slides were then incubated overnight at 4 °C with fluorophore-conjugated secondary antibodies (anti-mouse, anti-rat, anti-rabbit; 1:1,000 dilution) in SuperBlock (ThermoFisher Scientific) to assess non-specific binding and residual autofluorescence. Following washing (3 ×5 min PBS), slides were again photobleached using the same conditions.
For each CyCIF cycle, slides were incubated with Hoechst 33342 (1:10,000; Thermo Fisher Scientific) and either fluorophore-conjugated or unconjugated primary antibodies diluted in SuperBlock buffer (see Supplementary Table 4). Primary antibody incubation was performed overnight at 4 °C in the dark. Prior to imaging, slides were washed (3 ×5 min PBS), and an iSpacer (SunJin Labs; 0.15 mm, single-sided adhesive) was mounted around tissue borders to prevent compression. Slides were mounted in 70% glycerol, coverslipped with 24 ×50 mm No. 1.5 coverslips (Epredia), and imaged on a RareCyte CyteFinder II HT automated microscope.
Data processing and quality control
Cell segmentation was performed using the Xenium onboard pipeline in conjunction with the Xenium Cell Segmentation Kit. Nuclear masks generated from fluorescent nuclear staining were used as anchors for cell detection, and cytoplasmic boundaries were inferred by integrating nuclear positions with spatial transcript density derived from the 480-gene custom panel. RNA molecules were assigned to cells based on proximity and local density gradients, with geometric constraints applied to resolve overlapping cells. Segmentation outputs included per-cell polygon masks, centroids, and transcript counts.
Per-cell transcript counts were generated by aggregating molecule detections within segmented cell boundaries. Cells with low total transcript counts, elevated background signal, or poor segmentation metrics were excluded from downstream analyses. Quality-controlled data were imported into Python (v3.9 and v3.11) as AnnData objects, and expression values were analysed as raw counts or log-transformed values as indicated.
Xenium-CyCIF integration and alignment
Xenium and CyCIF images were aligned using the Image Registration functionality in Xenium Explorer (v4). Forty to fifty corresponding landmarks were manually selected between Xenium and CyCIF images to compute a landmark-based affine transformation. Following alignment, CyCIF cells were assigned Xenium cell identifiers based on centroid proximity, with an average centroid-to-centroid distance of 0.09 µm. The two datasets were integrated into a unified SpatialData object.
Downstream transcriptomic analysis
scVI batch correction was applied across individual tissue sections on the Xenium slide. Louvain clustering was performed on all scVI corrected transcriptomic data prior to cell-type annotation using canonical marker genes included in the custom panel. Cell-type annotations were manually refined using marker gene expression and spatial context. Differential gene expression analyses were conducted using non-parametric statistical tests with multiple-testing correction.
Downstream proteomic analysis
CyCIF data were processed to generate an AnnData object containing per-cell channel intensities for all imaged markers. Pro-Code knockout identities were assigned using the previously described delta-based intensity scoring deconvolution method.
Masson’s trichrome staining and analysis
Masson’s trichrome staining was carried out using Trichrome Stain Kit (Connective Tissue Stain, Abcam ab150686) on 5 µm-thick FFPE tissue sections following the manufacturer’s instructions.
For data analysis, tumour images were processed in Fiji. The Colour Deconvolution2 plugin with the ‘Masson trichrome’ vector was used to isolate collagen-specific staining. Area measurements were performed to quantify collagen content across all samples in a batch-wise manner.
DepMap analyses
For the visualization of DepMap data, we utilized the CRISPR (DepMap Public 25Q3+Score, Chronos) dataset, which was downloaded from the DepMap portal (https://depmap.org/portal/)68. This dataset included 46 human PDAC cell lines. In this dataset, a dependency score of 0 indicates that the gene is not essential for cell viability, while a score of −1 represents the median dependency score of universally essential genes, providing a benchmark for assessing gene essentiality.
Pan cancer bulk RNA-seq analysis
Data used for the pan cancer bulk RNA-seq analysis was generated by the TCGA (https://www.cancer.gov/tcga) and GTEx (https://www.gtexportal.org) projects. The raw data from all three projects were re-processed by Vivian et al., to reduced batch effects between the datasets69. In the same processing the data were log2-normalized. This re-processed data were accessed from the UCSC Xena Data Hubs via the UCSCXenaTools (v1.6.0) R package70. Specifically, the following datasets were used for the study: TcgaTargetGtex_RSEM_Hugo_norm_count, TcgaTargetGTEX_phenotype, TCGA_survival_data. Differences in normalized expression levels of SERPINE1 and SERPINB2 between cancer subtypes and normal tissue were tested using a two-sided Wilcoxon rank-sum test. Subsequently, P values were adjusted for multiple testing by the Benjamini–Hochberg correction. Patients with each cancer type were stratified into high- and low-expression groups for both genes using median gene expression as the threshold. To assess the survival advantage in the low-expression groups, the coxph function of the survival package (v3.7.0) was employed to calculate hazard ratios and p values. To analyse survival specifically in PDAC within the TCGA dataset, the data were subset, and Kaplan–Meier survival curves were generated using the survfit function of the survival package (v3.7.0), with P values calculated via log-rank tests.
Survival analysis in patients treated with immunotherapy
Bulk RNA-seq data and clinical data from Motzer et al. were analysed to assess the impact of SERPINE1 and SERPINB2 expression on progression-free survival (PFS) in patients treated with avelumab plus axitinib24. Patients were stratified into high- and low-expression groups for both genes using median gene expression as the threshold. Kaplan–Meier survival curves were fitted using the survfit function of the survival package (v3.7.0) and P values were computed with a log-rank test.
Survival meta-analysis
Bulk Transcriptomic data and clinical data from GSE7172948, E-MTAB-613449, GSE6245250, GSE22456452, GSE2873551, TCGA-PAAD (see above), ICGC-PACA-AU, and ICGC-PACA-CA (accessed through pdacR71 (v0.1.2)) were analysed to assess the effect of SERPINE1 and SERPINB2 on overall survival in patients with PDAC. Patients were stratified within each study into lowest and highest quartile groups. Then a Cox proportional hazards model was fit for every study using overall survival as the outcome and quartile-defined expression group as the predictor survival package (v3.7.0). Study-specific log(HR) and standard errors were then combined across using a random-effects meta-analysis in the meta package (v8.1-0) (REML τ² with Hartung-Knapp adjustment), visualized as a forest plot.
Mouse scRNA-seq
Fresh tumour samples were processed into single-cell suspensions following the “Flow Cytometry” tumour processing protocol described earlier. Cell viability was assessed using the Acridine Orange/Propidium Iodide viability staining reagent (Nexcelom). Suspensions with over 80% viability and minimal debris were deemed suitable for downstream experiments. scRNA-seq was conducted using the Chromium platform (10x Genomics) with the 5’ Gene Expression (5’ GEX) V2 kit, with a targeted recovery of 8,000 cells. Gel-Bead in Emulsions (GEMs) were generated on the sample chip K using the Chromium X (10x Genomics). Barcoded cDNA was extracted from GEMs through Post-GEM room temperature cleanup and amplified via 13 PCR cycles. Amplified cDNA was fragmented, end-repaired, poly A-tailed, adapter-ligated, and sample-indexed according to the manufacturer’s instructions. Libraries were quantified using TapeStation (Agilent) and QuBit (ThermoFisher) and sequenced in paired-end mode on an Illumina NovaSeq 6000 instrument, with a target depth of 25,000 reads per cell. Raw FASTQ files were aligned to the reference genome Gex-mm10-2020-A using CellRanger v5.0.1 (10x Genomics). Feature–barcode matrix files generated by CellRanger were utilized for downstream analyses.
For selected samples, dissected tumour tissues were immersion-fixed in 10% neutral buffered formalin for 24 h at room temperature and subsequently embedded in paraffin blocks for sectioning. FFPE tissue curls (20 µm thickness) were collected for RNA extraction using the RNeasy FFPE Kit (Qiagen) according to the manufacturer’s instructions. RNA quantity and integrity were assessed using an Agilent TapeStation with High Sensitivity RNA ScreenTape, and samples with DV200 values > 30% were considered suitable for downstream analysis. Single-cell RNA sequencing libraries were prepared using the Chromium Single Cell Gene Expression Flex protocol (10x Genomics), which is optimized for fixed and FFPE-derived RNA. In brief, paraffin was removed from tissue sections, and tissues were rehydrated and enzymatically digested to release single cells or nuclei following manufacturer’s instructions. Dissociated single cells or nuclei were incubated with transcriptome-wide probe sets, allowing hybridization to target RNA molecules retained within each cell or nucleus. Barcoded gel bead-in-emulsions (GEMs) were then generated on the Chromium X system, incorporating cell-specific barcodes and unique molecular identifiers (UMIs). Following GEM generation, libraries were amplified by PCR, subjected to sample indexing, and purified according to the manufacturer’s protocol. Final libraries were quantified using Qubit fluorometric analysis (Thermo Fisher Scientific) and Agilent TapeStation, pooled, and sequenced on an Illumina NovaSeq 6000 platform in paired-end mode, targeting 25,000 reads per cell. Raw FASTQ files were aligned to the reference genome Gex-mm10-2020-A using CellRanger multi v7.0.1 (10x Genomics). Feature–barcode matrix files generated by CellRanger were utilized for downstream analyses.
scRNA-seq analyses
scRNA-seq on fresh tissue
ScRNA-seq data were processed using the Scanpy library (v1.9.4). Data from the control, Serpinb2-KO and Serpine1-KO conditions were concatenated into a single AnnData object for joint analysis. Quality control included filtering cells with fewer than 800 or more than 40,000 total counts. Cells with mitochondrial gene content exceeding 20% of total counts were excluded to account for potential stress or degradation. Additionally, cells expressing fewer than 300 genes were removed. Doublets were identified using Scrublet (v0.2.3) with default parameters and filtered based on a threshold doublet score of 0.04. Genes with fewer than 20 total counts across all cells were removed. Raw counts were normalized using a log-transformation with a pseudocount of 1 to stabilize variance across cells. Highly variable genes were identified using Scanpy’s highly_variable_genes function, focusing on the top 4,000 genes based on dispersion and mean expression across cells. The processed dataset was used for subsequent normalization, dimensionality reduction, clustering, and annotation.
Chromium Single Cell FLEX
A matched workflow with additional steps to account for FFPE-specific technical noise was used for the analysis. Filtered feature–barcode matrices from control, Serpinb2-KO and Serpine1-KO samples were concatenated, and quality control thresholds analogous to those used for fresh scRNA-seq were applied, including filtering of cells with fewer than 800 or more than 40,000 total counts, mitochondrial gene content exceeding 20%, or fewer than 400 detected genes. Doublets were identified using Scrublet (v0.2.3) with default parameters and filtered based on a threshold score of 0.14. Ambient RNA contamination was corrected using scAR (v0.6.0). Batch integration across experimental conditions was performed using scVI (v1.3.3), trained on scAR-denoised counts with batch labels corresponding to condition. The resulting latent representations were used for neighbourhood graph construction, UMAP visualization, and Leiden clustering. Highly variable genes were identified in a batch-aware manner using Scanpy, selecting up to 2,000 genes for downstream analyses.
Differential gene expression analysis was performed with the tool rank_genes_groups, which is part of the Scanpy package. The Benjamini–Hochberg method was used to correct for multiple testing. Subsequent enrichment analyses were performed using the Enrichr database72, with cut-offs used for the specific analyses being specified in the text and corresponding figure legends.
The fibrinogen-binding score was calculated using the score_genes function in Scanpy and a curated gene set derived from the Gene Ontology term ‘fibrinogen binding’ (GO:0070051). The resulting per-cell scores reflect the relative expression of genes encoding proteins with known fibrinogen-binding activity and were used to assess enrichment of fibrinogen-associated programmes across macrophage subpopulations.
The additional dataset analysed can be found at GSE207938.
Human scRNA-seq analysis
Raw count matrices were obtained from GEO repositories of published studies. Datasets were processed individually for standard scRNA-seq pre-processing. In brief, doublets were removed using the Solo method, data were then processed in Scanpy (v1.9.8), with cells expressing fewer than 200 genes being filtered out. Further pre-processing involved filtering to remove cells where: (1) log1p_total_counts, log1p_n_genes_by_counts or the top 20 gene fraction exceeded five median absolute deviations; and (2) mitochondrial_counts percentage surpassed three median absolute deviations or more than 20. The subsequent data were normalized, log1p-transformed, pre-annotated and integrated using scANVI (v1.0.4). The latent representation was used to compute the nearest neighbours distance matrix for Leiden clustering. Cells were manually annotated using established single-cell atlases. InferCNVpy (v.0.6.1) was used to identify cancer cells.
To study SERPINE1/B2 expression heterogeneity in human tumours for each patient, we quantified SERPINE1 and SERPINB2 expression in malignant cells as (i) the mean expression among gene-positive cancer cells and (ii) the percentage of gene-positive cancer cells. We excluded samples with less than 100 cancer cells.
Differential expression analysis was carried out using scanpy.tl.rank_gene_groups. SERPINE1 and SERPINB2 cancer cells were compared to all other cancer cells with the Wilcoxon method, after removing mitochondrial and ribosomal genes and filtering out genes that are expressed in <5% of the cells. Significantly differentially expressed genes (P ≤ 0.05, log fold change ≥ 1) that were upregulated were used for the enrichment analysis. Enrichment analyses were performed using the Enrichr database72, with cut-offs used for the specific analyses being specified in the text and corresponding figure legends.
Human spatial transcriptomics
A total of 57 primary PDAC tumour samples were analysed10,33,34. Spatial transcriptomic data were processed in Python (v3.9 and v3.11) using Scanpy (v1.10.1). For each sample, expression matrices were stored as individual AnnData objects with associated metadata, including sample ID, patient ID, tissue of origin, and spatial information. Standard QC metrics (total_counts, n_genes_by_counts) were computed with sc.calculate.qc.metrics, and samples were concatenated at the dataset-level along the observation axis using sc.concat, with a library_id field tracking sample identity. Raw counts were preserved for subsequent SCVI and cell2location (v0.1.4) analysis.
Spatial transcriptomic deconvolution and cell-type specific gene expression
Spatial transcriptomic count matrices from the three PDAC cohorts10,33,34 were loaded as AnnData objects and harmonized. For each dataset, the raw count matrix was copied into layers[“counts”] to preserve counts for cell2location. Spot-level mitochondrial gene counts and fractions were computed, after which MT genes were removed from the expression matrices. Each spot was annotated with a dataset label (obs[“dataset”]) and a technology label (obs[“tech”]). A per-slide batch identifier for deconvolution (obs[“batch_c2l”]) was then constructed as a concatenation of dataset, technology, and slide/library ID. For subsequent cell2location analysis, genes were restricted to the intersection of shared features between the spatial object and the custom single-cell PDAC atlas created in this study (see ‘Human scRNA-seq’), leaving 16,102 features. Genes were further filtered with the cell2location filter_genes utility, with cell_count_cutoff = 10, cell_percentage_cutoff2 = 0.03, and nonz_mean_cutoff = 1.12, leaving 10,983 features. Fine-grained single-cell annotations were collapsed into a smaller and broader set of 26 cell types to avoid conflating the model. The reference model was subsequently trained using sample ID as a batch key, treatment as a categorical covariate, and annotated cell types. The model was trained with maximum epochs being 250, and a batch size of 2500. For deconvolution, the combined spatial object was set up with library_id as the batch key and passed, together with the learned signatures, to the cell2location model (fixed detection prior and N_cells_per_location = 10). The expected expression of each gene in each cell-type was then calculated using the posterior distribution as per the cell2location’s standard vignette. Gene expression of individual genes (SERPINE1, SERPINB2) was plotted in spatial coordinates using cell2location’s custom plot_genes_per_cell_type function, obtained from GitHub.
Calculation of tumour and stromal content and broad annotation of Visium spots
For each Visium spot, the cell2location posterior abundance matrix (expected number of cells per cell type per spot—for example, the q05 abundance summary) was converted to per-spot cell-type fractions. Cell types were then grouped into four biologically defined compartments: Cancer (‘Cancer cells’), CAFs (MyoCAF, iCAF, stromal fibroblasts), endothelial cells and ‘Other’ (acinar and islet cells, or residual signal when not specified). For each spot, the summed fractional abundance was computed for each group, yielding four compartment scores (cancer, CAF, endothelial and other). Thus, each spot was annotated at two resolutions: (1) a higher dimensional profile of cell2location-derived abundances for individual cell types; and (2) aggregated compartment scores that summarize tumour versus stromal composition and delineate tumour-dominant versus stroma-dominant regions. For the Pei cohort, which lacked matched single-cell data but provided high-confidence tumour-cell-enriched spot calls derived from integrated histopathology, spatial CNVs and tumour lineage markers, we used the reported fraction of tumour-cell-enriched spots per section as a slide-level prior on tumour coverage.
Analysis of SERPINE1+ and SERPINB2+ spots in tumour and stromal compartments
Post-cell2location normalization was performed in Scanpy. Library-size normalization and log-transformation were applied, followed by selection of highly variable genes using library_id as a batch covariate and computation of a PCA embedding. To correct for between-slide batch effects while preserving local structure, BBKNN (v1.5.1) (sc.external.pp.bbknn, batch_key = “library_id”) was used to construct a batch-corrected neighbourhood graph for downstream visualization and clustering. Broad spot annotations (Cancer, CAF, Endothelial, Other) were used to quantify the compartmental origin of SERPINE1 signal. For each sample, normalized SERPINE1 expression was summed across all spots within each compartment and divided by the total SERPINE1 expression in that sample, yielding the fraction of SERPINE1 expression attributable to Cancer, CAF, Endothelial, and Other compartments.
SERPINE1 and SERPINB2 spatial plots by compartment were generated using the broad spot annotations (Cancer, CAF, Endothelial, Other) defined above. For each section and gene, spot-level expression was visualized on the Visium grid with a per-section dynamic range set to the 99th percentile of non-zero expression values, ensuring comparable contrast across slides while avoiding saturation by a few outliers. To relate expression to compartment identity, spots were coloured by SERPINE1 or SERPINB2 intensity and overlaid with compartment-specific markings (Cancer, CAF, Endothelial, Other) on the same spatial coordinates. This provided a qualitative readout of where serpin signal was concentrated within tumour, fibroblast, endothelial, and other regions across all sections.
Spatial microenvironment analysis of SERPINE1+ and SERPINB2+ spots
To compare the local TME between SERPINE1/B2+ and SERPINE1/B2− tumour regions, cell2location posterior abundances (q05_cell_abundance_w_sf) were converted to non-cancer compositions per spot. Spots were labelled positive for the respective gene if the expression was >1, and negative if the expression was 0. Cancer cells and MyoCAFs, which made up the majority of each slide, were excluded, and remaining non-cancer cell types were summed, and each cell type was expressed as a fraction of the non-cancer total. Within each dataset, these non-cancer compositions were compared between serpin+ and serpin− cancer spots using two-sided Mann–Whitney tests per cell type, yielding effect sizes (difference in mean composition and Cohen’s d) and P values. Benjamini–Hochberg FDR correction was applied separately within each dataset across cell types.
For compartment-resolved analyses, SERPINE1+ and SERPINB2+ spots were further stratified by their previously annotated broad compartment labels (cancer, CAF and endothelial). Within each SERPINE1/B2+ compartment, cell2location abundances were renormalized over non-cancer, non-stromal reference types by excluding all stromal and epithelial compartments (cancer cells, MyoCAF, iCAF, stromal fibroblasts, endothelial cells, ductal, acinar, islet, Schwann). This produced immune-only composition profiles for SERPINE1/B2+ cancer, CAF and endothelial spots. Mean immune compositions were plotted as heat maps, and Mann–Whitney tests with FDR correction were used to identify immune populations enriched in SERPINE1/B2+ neighbourhoods in each compartment.
Statistical analysis
Statistical analysis was conducted using GraphPad Prism 10 or using the software indicated above. Details of statistical analysis are described in the corresponding sections of the methods and in the figure legends.
Materials availability
The nuclear Pro-Code library is available from Addgene (Plasmid Kit #1000000197). All other vector constructs used in the manuscript are available by request from the corresponding author.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.