10,000 Matching Annotations
  1. Last 7 days
    1. eLife Assessment

      This important study introduces a non-perturbative pulse-labeling strategy for yeast nuclear pore complexes (NPCs), employing a nanobody-based approach in order to selectively capture Nup84-containing complexes for imaging and biochemical analysis. The data convincingly demonstrate that a short induction period (20 minutes to 1 hour) yields a strong and sustained signal, enabling affinity purification that faithfully recapitulates the endogenous Nup84 interactome. This tool offers a powerful framework for investigating NPC dynamics and associated interactomes through both imaging and biochemical assays.

    2. Reviewer #1 (Public review):

      Summary:

      The authors present a nanobody-based pulse-labeling system to track yeast NPCs. Transient expression of a nanobody targeting Nup84 (fused to NeonGreen or an affinity tag) permits selective visualization and biochemical capture of NPCs. Short induction effectively labels NPCs, and the resulting purifications match those from conventional Nup84 tagging. Crucially, when induction is repressed, dilution of the labeled pool through successive cell cycles allows the visualization of "old" NPCs (and potentially individual NPCs) providing a powerful view of NPC lifespan and turnover without permanently modifying a core scaffold protein.

      Strengths:

      (1) A brief expression pulse labels NPCs, and subsequent repression allows dilution-based tracking of older (and possibly single) NPCs over multiple cell cycles.

      (2) The affinity-purified complexes closely match known Nup84-associated proteins, indicating specificity and supporting utility for proteomics.

      Weakness:

      Reliance on GAL induction introduces metabolic shifts (raffinose → galactose → glucose) that could subtly alter cell physiology or the kinetics of NPC assembly. As acknowledged by the authors, alternative induction systems (e.g., β-estradiol-responsive GAL4-ER-VP16) could be implemented as a way to avoid carbon-source changes.

      Comments on revised version.

      The authors have thoughtfully addressed all of my concerns. In particular, they have updated the proteomic analysis in Figure 1I, showing that they recover most NPC components (including basket Nups), including non-NPC proteins as controls, and providing all data as a supplementary table. These changes strengthen the authors conclusion and improve transparency. I have no further recommendations and congratulate the authors for their exciting work.

    3. Reviewer #2 (Public review):

      Summary:

      This preprint describes a practical and useful approach for labeling and tracking NPCs in situ, using a fluorescently conjugated nanobody that binds directly to the core scaffold nucleoporin Nup84 with nanomolar affinity. Useful applications including timelapse imaging, affinity purification, and proximity labeling are envisioned.

      Strengths:

      Clever use of a fluorescently conjugated nanobody that binds directly to the core scaffold nucleoporin Nup84 with nanomolar affinity.

    4. Reviewer #3 (Public review):

      Summary:

      Submitted to the Tools and Resources series, this study reports on the use of a single-domain antibody targeting the nucleoporin Nup84 to probe and track NPCs in budding yeast. The authors demonstrate their ability to rapidly label or pull down NPCs by inducing the expression of a tagged version of the nanobody (Fig. 1).

      Strengths:

      This tool's main strength is its versatility as an inexpensive, easy-to-set-up alternative to metabolic labelling or optical switching. This same rationale could, in principle, be applied to the study of other multiprotein complexes using similar strategies, provided that single-chain antibodies are available.

      Weaknesses:

      This approach has no inherent weaknesses, but it would be useful to verify in the future that this pulse labelling strategy can also be used to detect assembly intermediates, structural variants, or damaged NPCs, e.g. NPC clusters formed in some nucleoporin mutants.

      Overall, the data clearly shows that Nup84 nanobodies are a valuable tool for imaging NPC dynamics and investigating their interactomes through affinity purification.

      Comments on revised version.

      None at this stage.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors present a nanobody-based pulse-labeling system to track yeast NPCs. Transient expression of a nanobody targeting Nup84 (fused to NeonGreen or an affinity tag) permits selective visualization and biochemical capture of NPCs. Short induction effectively labels NPCs, and the resulting purifications match those from conventional Nup84 tagging. Crucially, when induction is repressed, dilution of the labeled pool through successive cell cycles allows the visualization of "old" NPCs (and potentially individual NPCs), providing a powerful view of NPC lifespan and turnover without permanently modifying a core scaffold protein.

      Strengths:

      (1) A brief expression pulse labels NPCs, and subsequent repression allows dilution-based tracking of older (and possibly single) NPCs over multiple cell cycles.

      (2) The affinity-purified complexes closely match known Nup84-associated proteins, indicating specificity and supporting utility for proteomics.

      We thank the reviewer for this evaluation

      Weaknesses:

      (1) Reliance on GAL induction introduces metabolic shifts (raffinose -> galactose -> glucose) that could subtly alter cell physiology or the kinetics of NPC assembly. Alternative induction systems (e.g., β-estradiol-responsive GAL4-ER-VP16) could be discussed as a way to avoid carbon-source changes.

      Indeed, this could be an improvement, and we mention the benefits of an inducible system that does not alter the cell’s metabolic state in the discussion on p.3.

      (2) While proteomics is solid, a comprehensive supplementary table listing all identified proteins (with enrichment and statistics) would enhance transparency.

      Indeed, we now provide source data showing LFQ intensities, fold-enrichment and statistics for all detected proteins.

      (3) Importantly, the authors note that the method is particularly useful "in conditions where direct tagging of Nup84 interferes with its function, while sub-stoichiometric nanobody binding does not." After this sentence, it would be valuable to add concrete examples, such as experiments examining NPC integrity in aging or stress conditions where epitope tags can exacerbate phenotypes. These examples will help readers identify situations in which this approach offers clear advantages.

      Indeed, we agree this would be useful. For example, in Nup1Δct and Nup60Δ mutants, GFP-tagging of Nup84 leads to slower growth and increased cell size (Ollivaud et al., BioRxiv). We have however not extensively tested nanobody expression in these mutants, and cannot conclude that it has no interfering effects. We therefore rephrased to “while sub-stoichiometric nanobody binding does may not, …”. Another situation where we find the nanobody-based labeling useful is when we want to assess the structural integrity (IPs) and localization (imaging) of NPCs in mutant strains, but prefer not to use tagged Nups in the actual experiments. In these cases, we transiently express the Nup84 nanobody to perform these checks, and then carry out the experiments without the nanobody to avoid any tag-related interference. We hence also added “,…or when the temporary introduction of a ZZ- or mNG-tagged nanobody allows assessment of the integrity or localization of mutant NPCs prior to performing experiments without the nanobody.

      We thank the reviewer again for the constructive feedback and thoughts.

      Reviewer #2 (Public review):

      Summary:

      This preprint describes a practical and useful approach for labeling and tracking NPCs in situ. While useful applications including timelapse imaging, affinity purification, or proximity labeling are envisioned, addressing some outstanding technical questions would give a clearer picture of the sensitivity and temporal resolution of this approach.

      Strengths:

      Clever use of a fluorescently conjugated nanobody that binds directly to the core scaffold nucleoporin Nup84 with nanomolar affinity.

      We thank the reviewer for this evaluation

      Weaknesses:

      The decrease in nanobody labeling over 8 hours of chase period is interpreted to indicate that NPCs turn over during this time. However, it is also possible that the nanobody: Nup84 association is disrupted during mitosis by phosphorylation, other PTMs, or structural remodeling.

      We thank the reviewer for this thought. It is actually not turnover that we propose to underly the decrease in nanobody labeling, but rather the dilution of labelled NPC to the daughter cell. The current data do not support the interpretation that the nanobody: Nup84 association is disrupted as proposed by the reviewer. The exchange of individual Nups, including Nup84, is slow with half-times in the order of hours (Hakhverdyan et al. 2021; Rabut, Doye, and Ellenberg 2004), and the nanobody: Nup84 association is very stable, namely in the nanomolar range (Nordeen et al. 2020). The association of nanobody with NPCs is thus expected to be very stable. Instead, dilution of labelled NPCs to the daughter - approximately 40% of the existing NPCs are transmitted to the daughter cell in each division (Zsok et al. 2024; Khmelinskii et al. 2010) – will lead to significant decreases in nanobody labelling over time. As the reviewer is likely aware, baker’s yeast NPCs – in contrast to mammalian NPCs - remain largely intact during cell division as there is no nuclear envelope breakdown.

      We thank the reviewer again for the constructive feedback and thoughts.

      Reviewer #3 (Public review):

      Summary:

      Submitted to the Tools and Resources series, this study reports on the use of a single-domain antibody targeting the nucleoporin Nup84 to probe and track NPCs in budding yeast. The authors demonstrate their ability to rapidly label or pull down NPCs by inducing the expression of a tagged version of the nanobody (Figure 1).

      Strengths:

      This tool's main strength is its versatility as an inexpensive, easy-to-set-up alternative to metabolic labelling or optical switching. This same rationale could, in principle, be applied to the study of other multiprotein complexes using similar strategies, provided that single-chain antibodies are available.

      We thank the reviewer for this evaluation

      Weaknesses:

      This approach has no inherent weaknesses, but it would be useful for the authors to verify that their pulse labelling strategy can also be used to detect assembly intermediates, structural variants, or damaged NPCs.

      We agree with the reviewer that it would be informative to see if VHH[Nup84] can bind its epitope in the context of an altered NPC structure but consider such studies to be beyond the scope of this study.

      Overall, the data clearly show that Nup84 nanobodies are a valuable tool for imaging NPC dynamics and investigating their interactomes through affinity purification.

      We thank the reviewer again for the constructive feedback and thoughts.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) In Figure 1A, and although it is partially mentioned in the legend, it would be helpful to indicate precisely when cells are grown in raffinose, when galactose is added for induction, and when glucose is used to terminate expression.

      We included “galactose” and “glucose” to Panel A to indicate induction and termination of expression, respectively.

      (2) Related to the previous point, consider mentioning the GAL4-ER-VP16 (ADGEV) estradiol-inducible system as an optional strategy to avoid carbon shifts and potentially reduce cell-to-cell variability.

      We mention the benefits of an inducible system that does not alter the cell’s metabolic state in the discussion on p.3

      (3) Add a brief sentence explaining that the ZZ tag is derived from Protein A and binds IgG Fc.

      This information is now added on p.2

      (4) The statement "all Nups significantly coenriched with VHH[Nup84]-ZZ..." is likely inaccurate, since not all Nups are labeled in panels F-H, and some basket components are missing in panel I (particularly basket components such as Nup60, Nup1). Consider revising to "most Nups significantly coenriched...". In panel I, please include a clearly non-enriched protein as a visual reference for the color scale.

      We are very grateful to the reviewer for pointing this out. We accidentally used a faulty filtering on the dataset to generate figure panel I, omitting several Nups that were reproducibly found in all replicas. All Nups, except for Gle1 and Pom33, were detected reproducibly.

      We have made the following adjustments to the figure panel and accompanying text:

      In Fig. 1I, we included the missing Nups and 5 proteins that co-purified with VHH[Nup84] but not specifically enriched, as the reviewer suggested. They cluster in a separate group and their abundance is not going up in time. We randomly selected these 5 proteins from the list of genes that were reproducibly found in all four timepoints.

      For clarity, we removed the NTRs

      We changed the text to “we found that all Nups, except Gle1 and Pom33, significantly coenriched with VHH[Nup84]-ZZ” on p.2.

      We updated the methods section, describing the clustering method and how we selected the 5 random proteins

      (5) Provide a supplementary spreadsheet with LFQ intensities, fold-enrichment, and statistics for all detected proteins. This will address questions about missing Nups and support transparency.

      This information is now added as Source data Figure 1.

      (6) Directly after the statement "Amongst others this is useful in conditions where direct tagging of Nup84 interferes with its function, while sub stoichiometric nanobody binding does not," it would be useful to include concrete instances, such as stress or aging conditions, where Nup84 tagging may sensitize NPC integrity.

      Indeed, we agree this would be useful. For example, in Nup1Δct and Nup60Δ mutants, GFP-tagging of Nup84 leads to slower growth and increased cell size (Ollivaud et al., BioRxiv). We have however not extensively tested nanobody expression in these mutants and cannot conclude that it has no interfering effects. We therefore rephrased to “while sub-stoichiometric nanobody binding does may not, …”. Another situation where we find the nanobody-based labeling useful is when we want to assess the structural integrity (IPs) and localization (imaging) of NPCs in mutant strains, but prefer not to use tagged Nups in the actual experiments. In these cases, we transiently express the Nup84 nanobody to perform these checks and then carry out the experiments without the nanobody to avoid any tag-related interference. We hence also added “,…or when the temporary introduction of a ZZ- or mNG-tagged nanobody allows assessment of the integrity or localization of mutant NPCs prior to performing experiments without the nanobody.

      (7) In panels K and L, since individual points correspond to biological replicates, overlaying a box plot obscures much of the data. Consider overlaying the means per replicate instead of box plots: see the "SuperPlots" approach for a clear explanation of how to present this (PMID: 32346721).

      We thank the reviewer for the “SuperPlots” suggestion, and we agree that representing the data in this way improves the visualization of individual biological replicates. We have updated the summarizing overlay in figures in panel K and L to represent the means per replicate instead of boxplots.

      (8) I spotted a few typos ("Lasty" ? "Lastly"; "in maintained" vs. "is maintained").

      Thank you, these are corrected

      Overall, this is a neat, well-executed methodological advance with clear value to the NPC field and potentially other complex assemblies. I look forward to seeing a revised version.

      Thank you!

      Reviewer #2 (Recommendations for the authors):

      Based on the recent structural analyses and NPC modeling using this nanobody, how accessible is the Nup84 epitope expected to be within the fully assembled NPC? While the data shown indicate that nanobody labeling of NPCs is readily detectable, stating this clearly would help motivate the approach and interpret the resulting data.

      We now included such a statement in the introduction on p.1.

      The decrease of nanobody labeling over 8 hours of chase period is interpreted to indicate that NPCs turn over due to cell division during this time window. However, it is also possible that nanobody:Nup84 association is disrupted during mitosis by phosphorylation, other PTMs, or structural remodeling.

      We thank the reviewer for this thought. It is actually not turnover that we propose to underly the decrease in nanobody labeling, but rather the dilution of labelled NPC to the daughter cell. The current data do not support the interpretation that the nanobody: Nup84 association is disrupted as proposed by the reviewer. The exchange of individual Nups, including Nup84, is slow with half-times in the order of hours (Hakhverdyan et al. 2021; Rabut, Doye, and Ellenberg 2004), and the nanobody: Nup84 association is very stable, namely in the nanomolar range (Nordeen et al. 2020). The association of nanobody with NPCs is thus expected to be very stable. Instead, dilution of labelled NPCs to the daughter - approximately 40% of the existing NPCs are transmitted to the daughter cell in each division (Zsok et al. 2024; Khmelinskii et al. 2010) – will lead to significant decreases in nanobody labelling over time. As the reviewer is likely aware, baker’s yeast NPCs – in contrast to mammalian NPCs - remain largely intact during cell division as there is no nuclear envelope breakdown.

      Reviewer #3 (Recommendations for the authors):

      (1) As mentioned above, to assess the general relevance of this tool, it would be informative to verify whether the VHH[Nup84] nanobody can access and detect NPC species under conditions that challenge their structural organization or biogenesis, for example, in nucleoporin mutants or under stress. The authors could, for instance, analyze the localization of VHH[Nup84] in yeast strains harboring clustered NPCs (nup133Δ), or following stresses known to impact NPC organization (e.g., osmotic stress or energy depletion; PMID: 34762489).

      We agree with the reviewer that it would be informative to see if VHH[Nup84] can bind its epitope in the context of an altered NPC structure and tried to include such data. Unfortunately, this was not successful, and further efforts are beyond the scope of his study. Following the reviewer’s suggestion, we expressed VHH[Nup84] in nup133∆N (nup133∆2-300) (Doye, Wepf, and Hurt 1994) following the experimental set-up in panel A and examined its localization. However, at t=2hrs hardly any nanobody signal was detectable in nup133∆N (see Author response image 1, upper panel A) and only after overnight expression nanobody-labelled NPC clusters are detectable (bottom panel A). Considering that expression levels of free mNG are also lower at t=2hrs in nup133∆N cells compared to WT cells (Author response image 1, panel B), it appears that protein expression under the Gal system is generally reduced in a nup133∆N background. These expression level differences between nup133∆N and WT preclude statements about the accessibility of the Nup84 epitope in nup133∆N. We note that nup133∆N cells do not have general mRNA export defects (Doye, Wepf, and Hurt 1994), so other inducible systems may be better suited for such analysis.

      Author response image 1.

      Expression level differences in WT and Nup133∆N cells. Left: localization of VHH[Nup84]-mNG in Nup133∆N cells at t=2hr following a 20-minute induction pulse and after overnight 0.5% galactose (ON) induction. Right: mNG levels in WT and Nup133∆N cells at t=2hr following a 20-minute induction pulse. Brightness/contrast settings are identical between the two panels. All panels are sum slices projections from 30 z-slices of 0.1µm. Scale bar = 5 µm.

      (2) Since outer rings are found on both sides of NPCs (i.e., the cytoplasmic and nuclear faces), could the authors indicate whether the VHH[Nup84] nanobody can enter the nucleus and probe the nuclear outer rings? Along these lines, it would be useful to provide a summary of the structural organization of NPCs in the introduction.

      Thank you, we have added a sentence on the localization of Nup84 in NPCs in the introduction. Based on what is known about influx (nuclear transport receptor-independent nuclear entry) of proteins with similar size and surface properties (Popken et al. 2015; Timney et al. 2016), the nanobody can rapidly enter the nucleus and hence bind Nup84 on both the nuclear and cytoplasmic side. We have no data to answer if binding might initially be biased towards cytosolic VHH[Nup84] binding the cytoplasmic outer rings.

      (3) The authors state that VHH[Nup84] and direct Nup84 detection are indistinguishable (p. 2). Could they provide images of the endogenously tagged Nup84-GFP strain for comparison?

      We have now included a pairwise comparison in a Figure 1 – supplement 1.

      Minor corrections:

      (1) There are a few typos that need correcting: 'Nup84Δ' (p. 1; should read 'nup84Δ') and 'promotor' (p. 2; should read 'promoter').

      Thank you, these are corrected

      (2) The reference 'Veldsink et al. 2025' (quoted in the PunctaFinder analysis description on page 8) does not appear in the References section.

      Thank you, these are corrected.

      We thank the reviewer again for the constructive feedback and thoughts.

      References

      Doye, V., R. Wepf, and E. C. Hurt. 1994. 'A novel nuclear pore protein Nup133p with distinct roles in poly(A)+ RNA transport and nuclear pore distribution', EMBO J, 13: 6062-75.

      Khmelinskii, Anton, Philipp J. Keller, Holger Lorenz, Elmar Schiebel, and Michael Knop. 2010. 'Segregation of yeast nuclear pores', Nature, 466: E1-E1.

      Popken, Petra, Ali Ghavami, Patrick R. Onck, Bert Poolman, and Liesbeth M. Veenhoff. 2015. 'Size-dependent leak of soluble and membrane proteins through the yeast nuclear pore complex', Molecular Biology of the Cell, 26: 1386-94.

      Timney, Benjamin L., Barak Raveh, Roxana Mironska, Jill M. Trivedi, Seung Joong Kim, Daniel Russel, Susan R. Wente, Andrej Sali, and Michael P. Rout. 2016. 'Simple rules for passive diffusion through the nuclear pore complex', Journal of Cell Biology, 215: 57-76.

      Zsok, J., F. Simon, G. Bayrak, L. Isaki, N. Kerff, Y. Kicheva, A. Wolstenholme, L. E. Weiss, and E. Dultz. 2024. 'Nuclear basket proteins regulate the distribution and mobility of nuclear pore complexes in budding yeast', Mol Biol Cell, 35: ar143.

    1. eLife Assessment

      This valuable study presents a comparative analysis of the transcriptomic features underlying C. elegans longevity, providing insights into how different changes in gene expression can promote longevity. The authors present solid evidence with analysis and selected functional validation showing that some long-lived animals share common changes while others appear to use opposing strategies. The datasets and analyses contained within and the user-friendly website developed will be of interest to researchers interested in complicated transcriptomic analyses and/or the biology of aging.

      [Editors' note: this paper was reviewed by Review Commons.]

    2. Reviewer #1 (Public review):

      This manuscript by Rudich ZD et al. systematically profiled the transcriptomic changes in nine long-lived C. elegans mutants and presented a careful and informative comparative analysis of these aging-related changes. In addition to these valuable datasets and bioinformatics analyses, the authors performed a large-scale RNAi screen to assess the role of the differentially expressed genes (DEGs) in these mutants and identify several potential targets to promote healthy aging. Moreover, the authors have provided a user-friendly website to examine genes of interest in those longevity mutants from their datasets.

      Strengths:

      Compared to previous transcriptomic analyses of these mutants in different reports, this study minimized the technical variations and benefitted from the advances in RNA-Seq technology and bioinformatics tools. Therefore, it should provide a more consistent and comprehensive view of the molecular mechanisms underlying the longevity of these mutants. The datasets in this manuscript are valuable to other researchers in the biology of aging.

      Weaknesses:

      Meanwhile, since these mutants have been extensively studied, the advance of this study in unknown ageing mechanisms remains limited.

      Comments on revised version.

      In the revised manuscript, the authors have addressed most of my concerns. In the text of this manuscript, the authors should still include more discussion on why osm-5 and daf-2 are categorized into two different groups.

    3. Reviewer #2 (Public review):

      Summary:

      In the manuscript titled "Multiple Molecular Pathways to Longevity: Opposing Gene Expression Programs Define Distinct Aging Strategies", the authors investigated diverse genetic pathways that contribute to lifespan extension in Caenorhabditis elegans and aimed to identify shared and distinct molecular mechanisms among various longevity mutants. Through comprehensive RNA sequencing of different longevity mutants representing seven distinct pathways, the authors showed that these mutants cluster into three primary groups based on their gene expression profiles. This transcriptomic analysis revealed that while some longevity genes are commonly regulated across multiple pathways, others exhibit opposing expression patterns, suggesting that distinct molecular strategies can lead to increased lifespan. Specifically, they identified a set of 196 genes that are consistently upregulated in most longevity mutants, many of which are involved in innate immunity and stress defense. By performing RNAi-based screening, the authors further validated the functional roles of several candidates, including C08F11.7, ugt-62, and K05C4.9, supporting their contributions to longevity and stress resistance. The authors conclude that longevity is mediated through multiple molecular pathways and provide a public online tool to study these complex transcriptomic landscapes.

      Significance:

      This study provides a systematic, side-by-side transcriptomic comparison of nine genetically distinct long-lived C. elegans mutants, revealing that lifespan extension arises from both shared and opposing gene expression programs. By identifying three distinct longevity groups and demonstrating that key pathways can be modulated in opposite directions to achieve long life, the work challenges the notion of a single universal transcriptional signature of aging. Importantly, functional validation shows that select commonly regulated genes can directly modulate lifespan and stress resistance, highlighting actionable molecular targets for promoting healthy aging.

      Comments on revised version:

      The authors addressed my concerns successfully.

    4. Author response:

      Reviewer #1:

      Major comments

      (1) Although I myself believe that the datasets in this study should be more consistent and comprehensive, the authors should perform a data mining analysis of previously reported transcriptomic changes of these mutants or similar mutants in the same longevity pathway and compare the reported changes with their findings to highlight the necessity and advances of this study.

      According to this suggestion, we have compared the differentially expressed genes identified in this study to previous gene expression studies involving these long-lived mutant strains. To our knowledge no previous studies have examined gene expression in sod-2 or ife-2 mutants, and at the time that we performed the RNA sequencing gene expression in osm-5 worms had not been examined (it took us a long time to complete this paper). We have included weighted Venn diagrams to illustrate the overlap and supplemental tables to list the overlapping di erentially expressed genes. For our current study, we felt it was important to compare RNA-seq data generated under exactly the same experimental and analysis paradigms in order to best compare across the nine long-lived mutants. These new analyses are included in Figures S19 – S25 and Table S2 . Please see lines 111-114, Figure S19-25, and Table S2.  

      (2) This manuscript does not perform any regulon or transcription factor (TF) analyses. TFs are the drivers of the transcriptomic changes and multiple conserved TFs (e.g., daf-16) have already been identified in these pathways. Therefore, it is necessary to examine and compare the regulons/TFs in these new datasets by bioinformatics. Such analyses can: a) provide more information of the driving force of these transcriptomic changes; b) show the role of these known longevity TFs; c) propose new TFs driving longevity; d) support the findings of 'longevity strategies' and 'longevity groups' from the perspective of TFs.

      According to this suggestion, we have now performed transcription factor analysis on the RNA-seq data to determine which transcription factors might be driving the longevity-associated transcriptional changes. To do this we used two complementary approaches: (1) transcription factor inference, which is based on the coordinated expression changes of known transcription factors; and (2) motif enrichment analysis, which is based on identifying transcription factor binding motifs in the promoters of di erentially expressed genes. After identifying which transcription factors were identified for each individual mutant, we then compared the identified transcription factors across all nine mutants. Interestingly, while 33 of the same transcription factors were implicated in group 1 and group 2 longevity mutants, 25 are modulated in different directions (activated in group 1, repressed in group 2 or vice versa) while only 5 are modulated in the same direction. This indicates that although group 1 and group 2 longevity mutants may modulate overlapping pathways to achieve long lifespan, in most cases these pathways are modulated in opposite directions. These new analyses are included in Figure S31 and Table S5. Please see lines 194-208, Figure S31, and Table S5.  

      (3) osm-5 and daf-2 are categorized into two different groups in this study. Since the longevity of cilia (-) mutants is through daf-16, the same master TF driving daf-2 longevity, please perform further analyses or discussion to clarify this issue.

      Loss of daf-16 is generally detrimental to lifespan. Disruption of daf-16 decreases the lifespan of all nine long-lived mutants that we examined (see supplemental table in our review paper PMID:37127095). However, loss of daf-16 also decreases wild-type lifespan. Thus, without further evidence it is hard to distinguish between the loss of daf-16 non-specifically decreasing lifespan verse activation of DAF-16 actually contributing to lifespan extension. In daf-2 mutants and the long-lived mitochondrial mutants there is increased nuclear localization of DAF-16 and upregulation of DAF-16 target genes. The differentially expressed genes in the long-lived mitochondrial mutants exhibit about a 50% overlap with the differentially expressed genes in daf-2 mutants (see Author response image 1). In contrast, osm-5 mutants show upregulation of some DAF-16 upregulated genes, no change in some DAF-16 upregulated genes and downregulation of other DAF-16 upregulated genes (see Author response image 1). Only about 10% of the differentially expressed genes in osm-5 mutants overlap with differentially expressed genes in daf-2 mutants. We believe that these results are consistent with loss of DAF-16 causing a general decrease in lifespan and not specifically contributing to osm-5 longevity. These comparisons will be included in a manuscript that we are currently preparing on osm-5 mutant longevity.

      Author response image 1.

      (4) This manuscript focused on genes whose RNAi suppressed the mutants longevity. Please also use bioinformatics to analyze the functions of those whose RNAi extends the mutants longevity, because these genes could tell the health price these mutants pay and help improve ageing interventions by reducing side effects.

      We perform enrichment analysis for both genes upregulated and downregulated in the long-lived mutant strains. The downregulated genes are involved in translation, ribosome biogenesis and gene expression. For the RNAi screen, we aimed to identify genes that are contributing to longevity and so we looked for a decrease in the lifespan of long-lived mutants when treated with RNAi. We did not screen for genes that extend the long-lived mutants longevity. While we did, nonetheless, identify multiple RNAi clones that increased either daf-2 or nuo-6 lifespan, there were not enough genes to identify any patterns of enrichment.

      (5) (OPTIONAL) I strongly suggest a comprehensive comparison of these transcriptomic changes in long-lived mutants with published age-related transcriptomic changes in wild type worms.

      According to this suggestion, we have now compared the differentially expressed genes that we identified in the nine long-lived mutants with genes that were found to be differentially expressed with aging. Interestingly, the group 2 long-lived mutants show a larger overlap for genes modulated in the opposite direction as aging (genes downregulated during aging are upregulated in eat-2 and osm-5 mutants). We have added this new analysis to our manuscript. Please see lines 210-223, Figure S32 and Table S6.

      Minor comments

      (1) Please further clarify the analysis of DEGs correlated with lifespan extension in Fig. 2 by a depiction. In Fig. 2C and D, please label data dots from different strains with different colors.

      According to this suggestion, each strain has been labelled a different colour.

      (2) In Fig. 3 and S20, please label the percentage of overlapping genes on top of each bars.

      We have now labelled the percentage of overlapping genes in Figure 3 and S20 (now S27).

      Reviewer #2:

      Major comments

      (1) While the authors identified a set of 196 upregulated genes, the rationale for narrowing these down to the three final candidates (C08F11.7, ugt-62, and K05C4.9) is not clearly described. The authors show that genetic inhibition of several genes, including DC2.5, C05B5.5, T07C4.5, and W03B1.7, decreases lifespan in both nuo-6 mutants and wild-type animals. However, the authors did not describe why these additional validated candidates, which also showed significant effects on longevity, were not pursued for further

      characterization. The authors should explicitly state the criteria used to prioritize these three genes over the other validated genes.

      Due to the costs and time involved in generating and characterizing new strains, we decided that we would select three strains to study further as a proof-of-principle. When deciding which genes to study further, we considered several approaches. In the end, we chose to use the strength/reproducibility of the increase in weighted mortality to identify genes with a clear, consistent impact. C05B5.5 and T07C4.5 were ruled out because they had an inconsistent impact on weighted mortality (Figure S28). W03B1.7 was ruled out because it did not have a strong enough e ect on weighted mortality (Figure S28). That narrowed it down to C08F11.7, ugt-62, DC2.5, and K05C4.9. Of those 4, C08F11.7, ugt-62, and K05C4.9 have the greatest consistent impact on weighted mortality (Figure S28) and so these genes were chosen. We have updated the manuscript to include this justification for focussing on C08F11.7, ugt-62, and K05C4.9. Please see lines 273-278.

      (2) The authors conclude that longevity can be mediated by multiple molecular pathways. However, it remains unclear whether these distinct strategies can operate simultaneously or are mutually exclusive. The authors need to test whether lifespan extension in a Group 1 mutant is further enhanced or suppressed by the knockdown of a key Group 2-specific genes. These experiments would help determine these pathways act additively, antagonistically, or as partially redundant survival programs.

      This is an excellent suggestion. While our data identify several genes that are regulated in opposite directions in group 1 and group 2 longevity mutants, we do not yet know the extent to which each of these genes contribute to the longevity of group 1 and group 2 mutants. The three genes that we focused on for further characterization (C08F11.7, ugt-62 and K05C4.9) are upregulated in group 1 longevity mutants but not group 2 mutants. Contrary to what might be expected, RNAi knockdown of these genes does not decrease the lifespan of the group 1 longevity mutant daf-2 but does decrease the lifespan of the group 2 longevity mutant eat-2. We recently reviewed the e ect of di erent resilience pathways on the lifespan of long-lived genetic mutants. Disruption of daf-16, sek-1, skn-1, hsf-1, ire-1 and trx-1 can decrease lifespan in both group 1 and group 2 longevity mutants, but also decreases lifespan in wild-type worms suggesting that at least in some mutants the e ect on longevity may be non-specific. Disruption of hif-1 does not a ect the longevity of group 2 mutants, but does a ect the lifespan of some group 1 mutants (clk-1, isp-1, nuo-6) but not others (daf-2, glp-1). To more definitively answer the question, it would be interesting to cross different combinations of group 1 and group 2 longevity mutants to see the extent to which different longevity groups synergize. This is something we are currently working on for a separate manuscript. We have added these points to the revised manuscript. Please see lines 363-381.

      (3) The authors provide interesting data on overexpression of the three candidate genes. However, whereas C08F11.7 clearly demonstrates both necessity and sufficiency for lifespan extension, overexpression of ugt-62 and K05C4.9 does not independently extend lifespan. To strengthen the manuscript, the authors should expand the discussion of these divergent results and clarify possible explanations.

      According to this suggestion, we have expanded our discussion to discuss possibilities of why these genes might be having different effects on lifespan. Please see lines 411-423.

      (4) Key citations are missing and the authors should add multiple citations including the following ones. Please cite the following paper and discuss the authors' finding with respect to the related work (Lee et al PMID: 40814218). Add citations in the sentence describing changes in the transcriptome of C. elegans associated with age (Lee et al., PMID: 38508494). Furthermore, please cite papers describing the overviews of survival assay using C. elegans (Kwon et al., PMID: 40436148, Hwang et al., PMID: 40436147).

      We have added the suggested citations to the revised manuscript. Please see lines 211 (Ref #39), 307 (Ref #41), 423 (Ref #54) and 436 (Ref #55).

      Minor comments

      (1) To improve readability, please provide the full names for all abbreviations at their first appearance in the manuscript.

      We have added the full names for each abbreviation on first appearance.

      (2) Please ensure that the labels in the figures match the text exactly. For instance, if different promoters are used for generating overexpression animals, it may be helpful to indicate the specific promoter in the figure panel or legend for clarity.

      We have ensured that the nomenclature in the text and the figures is the same. We have noted the promoter used for the overexpression strains in the figure legend.

      (3) For all lifespan and stress resistance assays, please include the total number of animals (n) and the number of independent biological replicates (N) in the figure legends to confirm statistical reliability.

      We have added the number of animals and independent biological replicates to the figures and figure legends.

      (4) Please clearly specify the exact developmental stage of the animals used for the survival assays in the Materials and Methods section.

      We have updated the methods to describe the developmental stages used for the survival assays.

    1. eLife Assessment

      This valuable study provides insights into the role of MATR3 in oocyte maturation and folliculogenesis, using conditional knockout mice and in vitro follicle culture systems to show that MATR3 is required for oocyte growth and gene transcription, with downstream effects on follicle development. The evidence is solid, but some minor inadequacies in replication of key methods and independent validation reduce confidence in the conclusions. The work will be of interest to researchers in reproductive biology and fertility.

    2. Reviewer #1 (Public review):

      Summary:

      This study aims to clarify MATR3's function and molecular mechanism in oocyte growth and maturation, explore its association with OMA and its potential as a diagnostic and therapeutic target using specific knockout mouse models, human OMA samples and multi-omics technologies. And it has fully achieved preset objectives with results strongly supporting conclusions. Specifically, it addresses the gap in the synergistic mechanism of epigenetic and secretory signals regulated by RNA-binding proteins (RBPs) in oocyte growth and enriches the molecular etiological spectrum of oocyte maturation disorders. It is the first time to reveal the conservative function of MATR3 in multiple species, providing a paradigm for cross-species research on RBPs in the field of reproductive biology. And it provides a new candidate target for OMA, a clinically refractory infertility disease, and is expected to promote the optimization of assisted reproductive technology and the development of precision medicine.

      Strengths:

      The strengths of this study are significant and prominent. First, the research system is comprehensive, integrating knockout mouse models, in vitro knockdown models, multi-species (mouse, porcine and human) verification, combined with scRNA-seq, LACE-seq, CO-IP and other multi-omics and molecular biology technologies, forming a complete and progressive evidence chain. Second, the mechanism analysis is in-depth, clarifying the dual molecular mechanisms of MATR3 regulating the transcriptional synthesis and secretion of GDF9 through "recruiting KDM3B to regulate H3K9me2 demethylation" and "directly binding to Rdx mRNA", with a clear logical closed loop. Third, the clinical correlation is close. It is the first time to find abnormal nuclear localization of MATR3 in oocytes of OMA patients, providing new clues for clinical disease mechanism research, and verifying the downstream function of GDF9 through rescue experiments, effectively enhancing the translational value of the results.

      Weaknesses:

      This study included only one OMA patient's oocyte sample. Without clinical screening for MATR3 mutations or abnormal expression, establishing a causal relationship between MATR3 and OMA remains difficult.

    3. Reviewer #2 (Public review):

      Summary:

      This study investigates the role of MATR3 in oocyte development and folliculogenesis using conditional knockout mouse models together with in vitro follicle culture and molecular analyses. The authors aim to determine whether MATR3 regulates oocyte maturation and follicle development and to explore potential mechanisms linking MATR3 function to transcriptional and epigenetic regulation in growing oocytes.

      Strengths:

      A major strength of the work is the use of a conditional knockout mouse model combined with complementary in vitro follicle culture approaches, which together provide a useful framework for examining gene function during oocyte development. The study also attempts to integrate cellular phenotypes with molecular analyses of transcriptional activity and epigenetic markers.

      Weaknesses:

      Several weaknesses limit the strength of the conclusions. These include insufficient validation of key experimental manipulations (such as the efficiency of MATR3 knockdown in siRNA experiments), limited quantification or statistical analysis for some datasets, inconsistencies between the text and presented data in certain figures, and incomplete methodological descriptions that make it difficult to fully evaluate reproducibility.

      Comments on revised version.

      Thank you for submitting the revised manuscript. I believe the revisions have substantially improved the quality and clarity of the study, and the authors have addressed the major concerns raised during the initial review.

    4. Reviewer #3 (Public review):

      Summary:

      The study aims to elucidate the dual molecular mechanisms of the RNA-binding protein MATR3 in oocyte growth and maturation. The authors propose that MATR3, highly expressed in growing oocytes (GOs), regulates oocyte quality through two pathways: epigenetically, by recruiting KDM3B to remove the repressive H3K9me2 mark at the Gdf9 locus to activate transcription; and post-transcriptionally, by binding Rdx mRNA to maintain microvillus structure for GDF9 secretion. This mechanism ensures oocyte-granulosa cell communication and female fertility. The study also explores the link between MATR3 and human oocyte maturation arrest (OMA).

      Strengths:

      The study proposes an innovative dual-mechanism model encompassing "epigenetic transcriptional activation and cytoskeletal regulation," which not only expands the functional understanding of RNA-binding proteins in chromatin regulation but also reveals the coordination between nuclear transcription and organelle structure. By integrating scRNA-seq and LACE-seq, the authors constructed a comprehensive regulatory network for MATR3, identifying both key targets and numerous potential molecules, thereby providing rich resources for future mechanistic studies. Furthermore, the inclusion of oocyte samples from human OMA patients directly links the basic findings to clinical reproductive disorders. Despite the limited sample size, this approach demonstrates strong translational potential.

      Weaknesses:

      The partial phenotypic improvement achieved by exogenous GDF9 supplementation suggests that the downstream effector pathways may involve a more complex network regulation, implying that the current interpretation of GDF9 central role could be further explored. Regarding the developmental abnormalities of granulosa cells in the conditional knockout model, their pathological origins require in-depth analysis to determine whether they represent primary alterations or secondary adaptive responses resulting from the loss of oocyte signaling.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study aims to clarify MATR3's function and molecular mechanism in oocyte growth and maturation, explore its association with OMA, and its potential as a diagnostic and therapeutic target using specific knockout mouse models, human OMA samples, and multi-omics technologies. And it has fully achieved preset objectives with results strongly supporting conclusions. Specifically, it addresses the gap in the synergistic mechanism of epigenetic and secretory signals regulated by RNA-binding proteins (RBPs) in oocyte growth and enriches the molecular etiological spectrum of oocyte maturation disorders. It is the first time the conservative function of MATR3 has been revealed in multiple species, providing a paradigm for cross-species research on RBPs in the field of reproductive biology. It also provides a new candidate target for OMA, a clinically refractory infertility disease, and is expected to promote the optimization of assisted reproductive technology and the development of precision medicine.

      Strengths:

      The strengths of this study are significant and prominent. First, the research system is comprehensive, integrating knockout mouse models, in vitro knockdown models, multi-species (mouse, porcine, and human) verification, combined with scRNA-seq, LACE-seq, CO-IP, and other multi-omics and molecular biology technologies, forming a complete and progressive evidence chain. Second, the mechanism analysis is in-depth, clarifying the dual molecular mechanisms of MATR3 regulating the transcriptional synthesis and secretion of GDF9 through "recruiting KDM3B to regulate H3K9me2 demethylation" and "directly binding to Rdx mRNA", with a clear logical closed loop. Third, the clinical correlation is close. It is the first time to find abnormal nuclear localization of MATR3 in oocytes of OMA patients, providing new clues for clinical disease mechanism research, and verifying the downstream function of GDF9 through rescue experiments, effectively enhancing the translational value of the results.

      Weaknesses:

      This study included only one OMA patient's oocyte sample. Without clinical screening for MATR3 mutations or abnormal expression, establishing a causal relationship between MATR3 and OMA remains difficult.

      We greatly appreciate positive comments and constructive feedback on our manuscript.

      We are encouraged that you recognize the novelty, rigour, and clinical relevance of our study on MATR3 in oocyte development and OMA. We have carefully considered your comments and revised the manuscript accordingly. We will further expand the OMA patient cohort in future studies to verify the causal relationship between MATR3 and OMA.

      Reviewer #2 (Public review):

      Summary:

      This study investigates the role of MATR3 in oocyte development and folliculogenesis using conditional knockout mouse models together with in vitro follicle culture and molecular analyses. The authors aim to determine whether MATR3 regulates oocyte maturation and follicle development and to explore potential mechanisms linking MATR3 function to transcriptional and epigenetic regulation in growing oocytes.

      Strengths:

      A major strength of the work is the use of a conditional knockout mouse model combined with complementary in vitro follicle culture approaches, which together provide a useful framework for examining gene function during oocyte development. The study also attempts to integrate cellular phenotypes with molecular analyses of transcriptional activity and epigenetic markers.

      Weaknesses:

      Several weaknesses limit the strength of the conclusions. These include insufficient validation of key experimental manipulations (such as the efficiency of MATR3 knockdown in siRNA experiments), limited quantification or statistical analysis for some datasets, inconsistencies between the text and presented data in certain figures, and incomplete methodological descriptions that make it difficult to fully evaluate reproducibility.

      We greatly appreciate your constructive comments and suggestions. We are grateful for the recognition of our conditional knockout mouse model and experimental design. We have carefully addressed all the weaknesses mentioned by the reviewer, including the validation of key experiments, quantitative and statistical analysis, consistency between text and figures, and detailed methodological descriptions. Details are described point-by-point below.

      Reviewer #3 (Public review):

      Summary:

      The study aims to elucidate the dual molecular mechanisms of the RNA-binding protein MATR3 in oocyte growth and maturation. The authors propose that MATR3, highly expressed in growing oocytes (GOs), regulates oocyte quality through two pathways: epigenetically, by recruiting KDM3B to remove the repressive H3K9me2 mark at the Gdf9 locus to activate transcription; and post-transcriptionally, by binding Rdx mRNA to maintain microvillus structure for GDF9 secretion. This mechanism ensures oocyte-granulosa cell communication and female fertility. The study also explores the link between MATR3 and human oocyte maturation arrest (OMA).

      Strengths:

      The study proposes an innovative dual-mechanism model encompassing "epigenetic transcriptional activation and cytoskeletal regulation," which not only expands the functional understanding of RNA-binding proteins in chromatin regulation but also reveals the coordination between nuclear transcription and organelle structure. By integrating scRNA-seq and LACE-seq, the authors constructed a comprehensive regulatory network for MATR3, identifying both key targets and numerous potential molecules, thereby providing rich resources for future mechanistic studies. Furthermore, the inclusion of oocyte samples from human OMA patients directly links the basic findings to clinical reproductive disorders. Despite the limited sample size, this approach demonstrates strong translational potential.

      Weaknesses:

      The partial phenotypic improvement achieved by exogenous GDF9 supplementation suggests that the downstream effector pathways may involve a more complex network regulation, implying that the current interpretation of GDF9's central role could be further explored. Regarding the developmental abnormalities of granulosa cells in the conditional knockout model, their pathological origins require in-depth analysis to determine whether they represent primary alterations or secondary adaptive responses resulting from the loss of oocyte signaling.

      We greatly appreciate your positive and insightful comments on our study. We are grateful for the recognition of our novel dual-mechanism model, comprehensive multi-omics analysis, and translational potential from basic research to clinical OMA. We have carefully addressed the weaknesses raised by the reviewer, including in-depth discussion of the GDF9-centered regulatory network and clarification of the origin of granulosa cell abnormalities. More details point-by-point responses are provided below.

      Recommendations for the authors:

      Point-by-point responses to reviewers’ comments

      We thank the reviewer very much for his/her reviewing of our work, and we appreciate the constructive comments and suggestions that have helped us to prepare an improved revision. Based on the comments of the reviewer, we have carefully revised the manuscript by performing some new experiments.

      Reviewer #1 (Recommendations for the authors):

      (1) Did most of the follicles cultured in vitro reach the antral follicle stage after 6 days?

      We greatly appreciate your insightful question. We statistically analyzed the survival rate and antral follicle ratio of in vitro-cultured follicles after 6 days of culture. Due to differences in culture systems and protocols, the follicle survival rate in our study (57.43 ± 3.11%) was different from that reported in previous literature (92 ± 10%). However, the proportion of antral follicles among surviving follicles was highly consistent between our results and published data (83 ± 13% vs 80.87 ± 3.27%) (Cortvrindt and Smitz 2002).

      Author response image 1.

      Ratio and survival rate of antral follicles after 6 days of culture. n = 3. Data are represented as mean ± SD.

      (2) In Figure 2F, at which stage did MATR3 begin to affect oocyte diameter?

      Thank you for your careful observation. Our morphological analysis of oocytes collected from PD14 and PD23 mice showed no significant difference in oocyte diameter between the cKO and Ctrl groups at the GO stage (Fig. S3D, E). However, oocytes in the cKO group became significantly smaller than those in the Ctrl group once they reached the FGO stage (Fig. 2E, F). Taken together, these results indicate that the growth defect caused by MATR3 deletion begins to manifest during the transition from the GO to FGO stage, with significant reduction in oocyte diameter clearly observed at the FGO stage as shown in Figure 2F.

      (3) What was the developmental potential of oocytes in Matr3-knockout mice?

      Thank you for this important question. Compared with the Ctrl group, oocytes derived from cKO mice showed a drastically reduced fertilization rate (91.55 ± 1.96% vs 10.55 ± 4.78%) and almost completely failed to develop to the blastocyst stage (76.62 ± 7.56% vs 3.67 ± 3.38%). These results clearly demonstrate that maternal deletion of Matr3 severely compromises the developmental potential of mouse oocytes, including fertilization capacity and subsequent early embryonic development.

      Author response image 2.

      Results of in vitro fertilization of oocytes. 2-cell: 2 days after fertilization; blastocyst: 4 days after fertilization. Data are represented as mean ± SD. ***P < 0.001.

      (4) The legend labels in the figures should not be bold.

      Thank you for your valuable suggestion. We have revised all the figures accordingly.

      Reviewer #2 (Recommendations for the authors):

      This manuscript investigates the role of MATR3 in oocyte development and folliculogenesis using conditional knockout (cKO) mouse models combined with in vitro follicle culture approaches. The topic is relevant to the field of reproductive biology and provides potentially important insights into the molecular mechanisms regulating oocyte maturation and follicle development.

      While the study presents interesting observations and utilizes both in vivo and in vitro experimental systems, several issues need to be addressed before the manuscript can meet the expected standards. These include concerns related to data interpretation, validation of experimental approaches, completeness of methodological descriptions, and clarity in data presentation. In addition, the manuscript requires substantial language editing to improve clarity and readability.

      The comments below outline major issues that should be addressed to strengthen the manuscript, as well as specific minor points regarding presentation and clarity.

      (1) The manuscript requires substantial revision to improve the written language and grammar. Numerous sentences are unclear or awkwardly phrased, which makes interpretation of the results difficult in several sections. The authors are strongly encouraged to have the manuscript professionally edited or thoroughly revised for language and clarity before making a resubmission.

      We sincerely appreciate the careful and constructive comments on the language quality and clarity of the manuscript. We fully agree that the written language, grammar, and sentence structure need substantial improvement to ensure the results are presented clearly and accurately.

      To address these concerns thoroughly, we have carefully revised the entire manuscript, including correcting grammatical errors, refining awkward phrasing, and restructuring unclear sentences to enhance readability and logical flow. In addition, we have sought professional language editing support to further polish the English expression and ensure the manuscript meets the linguistic standards of the journal.

      All revisions related to language and clarity have been completed, and we believe the revised version is significantly improved in terms of readability and precision.

      (2) Interpretation of oocyte maturation results (Line 140; Figure 2E, H). The manuscript states: "During in vitro maturation, oocytes isolated from PD23 cKO mice could not develop to metaphase II (Fig. 2E, H)." However, Figure 2H appears to show that a small proportion of knockout oocytes do reach the MII stage. Therefore, the description in the text seems inconsistent with the data presented. The authors should clarify the exact maturation rates in both groups, revise the text to accurately reflect the data, and provide statistical analysis to support the stated conclusions.

      Thank you for your valuable comment. A small proportion of knockout oocytes from PD23 cKO mice can indeed develop to the MII stage. We have revised the corresponding description and supplemented the statistical analysis of maturation rates to support our conclusion.

      Line 140: “During in vitro maturation, oocytes isolated from PD23 cKO mice could not develop to metaphase II (Fig.2E, H).” have been replaced by “During in vitro maturation, the proportion of oocytes from PD23 cKO mice developing to metaphase II stage was significantly reduced (Fig.2E, H, 54.9±2.08% vs 9.57±1.11%).”

      (3) Human oocyte sample size: In Figure 1D, it is unclear how many human oocytes were analyzed. It is important to specify the sample size (n) for all experiments. The authors should clearly indicate the number of oocytes analyzed in this experiment. Provide this information either in the figure legend or in the main text.

      Thank you for this important comment. We agree that the altered subcellular localization of MATR3 in human OMA oocytes is of great physiological significance for understanding the functional role of MATR3 during oocyte development.

      Unfortunately, during a 3‑month period of sample collection, we examined MATR3 localization in immature oocytes that failed to reach the MII stage, obtained from 11 women undergoing IVF treatment. Among these samples, only one donor’s oocytes exhibited the NSN chromatin configuration. Excitingly, these NSN‑stage oocytes from this donor clearly showed the loss of MATR3 nuclear localization, which strongly supports the critical role of MATR3 during oocyte growth and maturation. We have now clearly stated the sample size (n = 11) in the figure legend and main text as suggested. In future studies, we will continue to collect more human oocyte samples to further validate these observations with an expanded sample size.

      (4) Figure annotation issue: The figure legend for Figure 1F refers to an arrow, but no arrow is visible in the figure panel. Please correct this inconsistency by either adding the appropriate arrow to the figure or revising the legend accordingly.

      Thank you for pointing out this error. We have revised the figure legend for Figure 1F accordingly to correct this inconsistency.

      (5) Description of follicle analysis (Line 146): The sentence: "This was reinforced by the data of available follicles within the follicles of mice on PD35 (Fig. 2I, J)." is incorrect or poorly phrased. It should likely read: "...available follicles within the ovaries of mice at PD35...".

      We really appreciate your constructive suggestion on the phrasing. We have revised this sentence in the revised manuscript accordingly.

      (6) Quantification of proliferating cells: Figure 2K shows Ki-positive cells, but quantitative analysis is not provided. The authors should quantify the number or proportion of Ki-positive cells in both control and cKO groups and include statistical analysis to support any claims regarding differences in proliferation.

      Thank you for your valuable suggestion. We have quantified the number of Ki‑67‑positive granulosa cells in both control and cKO groups and performed the corresponding statistical analysis.

      The quantitative results have been added to Fig. S3G, and the relevant description has been supplemented in the main text at Line 150 to support our conclusion regarding cell proliferation differences.

      Line 150: “Consistently, immunofluorescence staining showed that the numbers of Ki67-positive (Fig. 2K) in cKO mice were lower than those found in the Ctrl.” have been replaced by “Consistently, immunofluorescence staining showed that the numbers of Ki67-positive (Fig. 2K, Fig. S3G) in cKO mice were lower than those found in the Ctrl (73.55±13.29% vs 24.65±7.80%).”

      (7) Validation of findings in the in vivo cKO model (Figure 3): The development of an in vitro follicle culture system is an interesting and valuable component of the study. However, several key analyses performed in vitro (e.g., transcription assays and analysis of epigenetic markers) should ideally also be validated in oocytes derived from the in vivo cKO model.

      Thanks for the valuable concern. We fully agree with you that the in vivo cKO model should be used to validate several key analyses performed in vitro. We collected growing oocytes from Ctrl and cKO mice and conducted transcription assays as well as analysis of epigenetic markers. The results showed that Matr3 knockout significantly downregulated transcriptional activity in GO and increased H3K9me2 levels (Author response image 3), which is consistent with our in vitro findings (Fig 3D E I J). These in vivo results confirm that MATR3 plays a critical role in regulating GO transcriptional activity and H3K9me2 levels.

      Author response image 3.

      Matr3 knockout results in the reduction of transcriptional activity. A EU staining (green) in GO collected from Ctrl and cKO. n = 15. B Quantification of the mean fluorescence intensity of EU in oocytes. C H3K9me2 staining (red) in GO collected from Ctrl and cKO. n = 15. D Quantification of the mean fluorescence intensity of H3K9me2 in oocytes. Scale bar: 20 μm. Data are represented as mean ± S.D. ***P < 0.001.

      (8) To strengthen the conclusions, the authors should consider repeating key experiments using oocytes directly isolated from the cKO mice. This would help confirm that the observed effects are not artifacts of the in vitro culture system.

      Thank you for this valuable and constructive suggestion. We fully agree that the conditional knockout mouse model is essential for verifying the physiological significance of MATR3 in vivo.

      To address this point, we have validated multiple key in vitro findings using oocytes directly isolated from cKO mice. For instances, the changes in oocyte transcriptional activity (EU staining) (in Comments 7), H3K9me2 levels (in Comments 7), GDF9 levels (Fig 4.B C E F), and OO-Mvi (Fig 6.A B, Author response image 4) all showed consistent trends with our in vitro knockdown results. In addition, the complete infertility phenotype of cKO female mice further demonstrates that MATR3 is indispensable for oocyte growth and meiotic maturation. We have also provided supplemental data from GDF9 rescue experiments and sequencing analysis performed in the mouse model.

      Author response image 4.

      Matr3 knockdown impairs the structural integrity of oocyte OO-MVi. A p-ERM staining (green) showing the OO-Mvi in oocyte from NC and si-Matr3. B Quantification of the number of Oo-Mvi vesicles (n = 6). Scale bar: 20 μm. Data are represented as mean ± SD. ***P < 0.001.

      In conclusion, the core conclusions of this study are supported by the mutual validation of key experimental results from MATR3-specific knockdown in vitro and Matr3 conditional knockout mouse models in vivo.

      (9) (1) Validation of MATR3 knockdown: The in vitro MATR3 knockdown experiment presented in Figure 4G raises an important concern: it is unclear whether Matr3 knockdown was effectively achieved in the oocytes analyzed. The authors should provide direct evidence of knockdown efficiency, for example, immunostaining for MATR3 protein on the oocytes. Without such validation, it is difficult to interpret the functional outcomes observed.

      As requested, we have provided direct evidences of the knockdown efficiency via immunostaining, which is now presented in Fig.S2D.

      In this experiment, oocytes from early growing follicles (approximately 150 μm in diameter) were microinjected with Matr3 siRNA. Following 5 days of continuous in vitro culture, oocytes from both the NC and si-Matr3 groups were isolated and subjected to immunofluorescence staining to assess protein levels. As shown in the figure, the oocytes at this stage exhibited the characteristic non-surrounded nucleolus (NSN) chromatin configuration. We observed robust MATR3 protein expression within the nucleus of NC oocytes, whereas the MATR3 protein levels were markedly reduced in the si-Matr3 group. These results confirm the successful construction of the Matr3 knockdown model in early growing follicle oocytes.

      (9) (2) Furthermore, it would be more convincing if the authors could perform the Gdf9 supplementation experiments using follicles isolated from the cKO mice, rather than relying solely on siRNA knockdown in vitro. Such experiments would provide clearer and more physiologically relevant evidence. If these experiments were attempted but did not produce similar results, this should be discussed.

      Thank you for your valuable and insightful suggestion. We fully agree that performing GDF9 supplementation experiments using follicles isolated from cKO mice would provide more direct and physiologically relevant evidence to strengthen our conclusions.

      Unfortunately, when we attempted to conduct GDF9 rescue experiments on follicles from cKO mice, neither the control nor cKO follicles were able to develop to the antral follicle stage (n=3). We speculate that this was caused by insufficient bioactivity of the veterinary-grade FSH been used, as compared to the imported FSH been provided by NHPP. Unfortunately, this particular FSH product has been discontinued. We are currently actively seeking and attempting to purchase new, qualified FSH reagents to repeat these experiments and further validate our findings in future work.

      (10) Figure citation order: Figures are not cited sequentially in the text. For example, Figure 6J is described first (line 250), followed by Figure 6A. Figures should be discussed in logical order, typically starting from panel A. Please revise the text to ensure that figure panels are introduced sequentially.

      Thank you for this careful and important comment.

      We have carefully revised the citation order of all figure panels in the main text, especially for Figure 6, to ensure they are introduced sequentially from panel A to the last panel in logical and numerical order, rather than being cited out of sequence.

      The corresponding adjustments have been made in the revised version of the manuscript.

      (11) Figure 6J interpretation: The purpose of the images shown in this figure is unclear. The authors should provide higher magnification images to clearly visualize the Oo-Mvi structures and include quantification of the observed phenotype to support the interpretation. It is important because the main findings of the paper heavily rely on these results.

      Thank you for this valuable and constructive suggestion. We fully agree that higher‑magnification images and quantitative analysis are essential to clearly demonstrate the Oo‑Mvi structures and reliably support our conclusions, especially given the importance of these results to the main findings of this study.

      Accordingly, we have replaced the original panels in Figure 6A with higher‑magnification images to better visualize Oo‑Mvi structures. In addition, we have supplemented the corresponding quantitative analysis of the observed phenotype to strengthen the interpretation of this figure (Fig 6B). All revisions have been incorporated into the revised manuscript.

      (12) Incomplete Materials and Methods section: The Materials and Methods section lacks important experimental details required for reproducibility. Specifically, the Matr3 flox mouse model. Either provide the appropriate reference describing the Matr3 floxed mice or include details on how the floxed allele was generated.

      Thank you for your valuable and careful comment. We fully agree that detailed experimental information in the Materials and Methods section is crucial for ensuring the reproducibility of the study. And we apologize for the omission of key details regarding the Matr3 flox mouse model.

      In response to your suggestion, we have thoroughly supplemented the relevant experimental details in the Materials and Methods section of the revised manuscript, including the specific construction strategy of the Matr3 floxed allele. These detailed descriptions will enable other researchers to reproduce our mouse model and verify the experimental results.

      All supplementary information has been integrated into the revised manuscript to meet the requirements of experimental reproducibility. We greatly appreciate your guidance in helping us improve the completeness and rigour of our study.

      (13) Follicle isolation: The manuscript does not describe how growing follicles were isolated. Please specify whether follicles were isolated using enzymatic digestion or mechanical dissection and provide sufficient methodological detail so that other researchers can reproduce the experiments.

      Thank you for this valuable comment. We agree that adding this information is essential for ensuring the reproducibility of our experiments. We have supplemented the corresponding description in the Materials and Methods section. Briefly, growing follicles were isolated by mechanical dissection using insulin syringes under a stereomicroscope, without any enzymatic digestion.

      Reviewer #3 (Recommendations for the authors):

      (1) Since KDM3B and MATR3 interact in cell lines, does this relationship affect the functional localization of KDM3B within oocytes? Specifically, does the localization of KDM3B change in cKO mice (e.g., nuclear export or aggregation)?

      Thank you for this insightful and constructive question. To address whether the interaction between KDM3B and MATR3 influences the functional localization of KDM3B in oocytes, we performed immunofluorescence staining to examine the subcellular distribution of KDM3B in cKO oocytes. Our results demonstrated that the nuclear localization of KDM3B remained unaltered; no obvious nuclear export or abnormal aggregation was observed in MATR3-deficient oocytes (Author response image 5).

      Based on these observations combined with our other experimental data, we propose that MATR3 regulates oocyte transcriptional activity through its physical interaction with KDM3B, rather than by controlling the nuclear targeting of KDM3B. Notably, despite unchanged nuclear localization of KDM3B in MATR3 cKO oocytes, we detected significantly elevated global levels of H3K9me2 and markedly reduced transcriptional activity (in Comments 7). These findings indicate that KDM3B loses its physiological function of demethylating H3K9me2 and promoting transcription in the absence of MATR3.

      In line with this mechanism, previous results showed that KDM3B knockout in female mice leads to follicle arrest at the secondary follicle stage and consequent infertility (Liu et al. 2015). Collectively, we conclude that in MATR3 cKO oocytes, although KDM3B is properly localized in the nucleus, it fails to execute its H3K9me2 demethylase activity, thereby impairing normal transcriptional regulation during oocyte development.

      Author response image 5.

      Matr3 knockout has no effect on KDM3B localization. KDM3B staining (red) in GO collected from Ctrl and cKO. n   = 50. Scale bar: 40 μm.

      (2) As the GDF9 rescue experiment only partially restores the phenotype, it is suggested to select 2-3 novel targets with high binding intensity and significant expression changes from the LACE-seq data (e.g., Igf2bp2 or Ccnb1 mentioned in the text) to further illustrate that MATR3 regulates a network.

      Thanks for this meaningful and constructive suggestion.

      We have supplemented the relevant data in Fig. S9 and further elaborated on these findings in the Discussion section in our revised manuscript, following your advice. Briefly, we collected growing oocytes from Ctrl and cKO mice and performed RT‑qPCR analysis to verify the expression of Igf2bp2 and Ccnb1 - two representative novel targets with strong binding intensity and significant expression changes identified from our LACE‑seq bioinformatics analysis. The results showed that both genes were significantly downregulated in cKO oocytes compared with controls, supporting the notion that MATR3 regulates a functional RNA network during oocyte development.

      (3) Are the granulosa cell defects primary or secondary? It is recommended to collect ovaries from earlier-stage cKO mice (e.g., PD7 or PD10) to examine the levels of FOXL2 and PCNA in granulosa cells.

      Thank you for your valuable comment. To clarify whether the granulosa cell defects are primary or secondary, we further investigated the temporal effect of MATR3 deficiency on granulosa cells by collecting ovaries from PD7, which is a critical period for primordial follicle activation and early follicular development.

      To evaluate the status of granulosa cells, we performed immunofluorescence staining on PD7 ovarian sections using FOXL2 (a specific marker for granulosa cells) to quantify the number of granulosa cells, and Ki67 (a proliferation-related marker) to assess the proliferative capacity of granulosa cells. The results, as shown in Author response image 6, demonstrated that there were no significant differences in either the number of granulosa cells or their proliferation levels in primary follicles between cKO and Ctrl.

      These findings are consistent with the data in our supplementary Fig S3F, where we observed no significant differences in the number of primordial follicles and growing follicles at PD7 between the two groups. Collectively, these results indicate that the activation of primordial follicles is not affected by MATR3 deficiency in oocytes, and the impairment of granulosa cells caused by oocyte-specific Matr3 knockout occurs at the secondary follicle stage rather than the early stages of follicular development. Therefore, we conclude that the granulosa cell defects in cKO mice are secondary to the oocyte dysfunction induced by MATR3 deficiency.

      Author response image 6.

      Loss of MATR3 in oocytes does not affect the number and proliferation of granulosa cells in primordial follicles. A Immunohistochemistry results showing granulosa cells in PD7 ovaries from Ctrl and cKO. B Quantification of granulosa cell number in the largest cross-section of primary follicles. C Quantification of the proliferation rate of granulosa cells in the largest cross-section of primary follicles. n = 15. Data are represented as mean ± SD. n.s., not significant.

      References:

      (1) Cortvrindt RG, Smitz JE. 2002. Follicle culture in reproductive toxicology: a tool for in-vitro testing of ovarian function? Human reproduction update 8: 243-254.

      (2) Liu Z, Chen X, Zhou S, Liao L, Jiang R, Xu J. 2015. The histone H3K9 demethylase Kdm3b is required for somatic growth and female reproductive function. International journal of biological sciences 11: 494-507.

    1. eLife Assessment

      This study reports important and invaluable findings that advance understanding of how attention is distributed between what we look at directly and what lies outside the center of gaze during active visual search. The evidence supporting the main claims is solid, with a large and rich dataset spanning multiple brain areas, although some aspects of the interpretation would benefit from additional controls and clearer separation of attention from eye-movement planning. The work will be of particular interest to researchers studying attention, visual perception, and eye movements behavior.

    2. Reviewer #2 (Public review):

      Summary:

      In natural visual behavior, such as when one is looking for a face in the crowd, the eyes are moved from site to site, seeking possible matching targets. This involves attention both to the current view at center of vision (the foveal location) as well as to upcoming views via attention to targets in the periphery. While it has been established that attention generally enhances neuronal response (compared to simple visual activation) at the attended spatial location, this study provides solid evidence that attention during active visual search leads to neuronal response enhancement only when the eye moves towards targets that exhibit the desired feature and category. This study thus moves the field towards understanding the neural encoding of active vision.

      This study examines the neuronal basis of feature selective attention during active, freely behaving visual search. Traditional electrophysiological studies on visual attention in monkeys commonly used an eye fixation with covert attention paradigm, but have not sufficiently addressed the roles of both foveal and peripheral attention in play during natural looking behavior. Here, the authors present a novel paradigm in which, during eye movement mediated search neuronal receptive fields are recorded in multiple cortical areas (sensory V4, temporal and prefrontal areas). In this manner, as the eye foveates, items in the array fall into foveal or non-foveal recorded sites. Thus, the experimental paradigm is elegant, offering the opportunity to make multiple types of comparisons: target/distractor, towards/away from fovea, areal. Specifically, following a category cue (face, house, hand, flower), freely initiated saccades are made to locate a categorically matching 'target' in an array of distractors. Feature attention is assessed by comparing eye saccades made to targets vs to distractors. Spatial attention is assessed by comparing saccades made 'towards' vs 'away' from targets. Statistics are rigorous and nicely designed. Detailed association of simultaneously obtained eye movement sequences and neural parameters are well done. These are valuable data which will contribute to our understanding of attentional modulation in visual search.

      The significance of these findings is fundamental. Decades of attention research in vision have been based on the paradigm of visual fixation and covert peripheral attention. However, increasingly the field has moved towards understanding how the visual system works during active vision. Here, the authors use an active visual search paradigm and record from key mid-tier (V4) and higher order (IT, PFC) areas. They find enhancement of attention both in the foveal and peripheral locations, and, furthermore, marked by a high degree of feature and categorical specificity. That is, while attention generally enhances neuronal response (compared to simple visual activation) at the attended spatial location, this study provides solid evidence that attention during active visual search leads to neuronal response enhancement only when the eye moves towards targets that exhibit the desired feature and category. This provides valuable data for the concept of a foveal-peripheral spatiotemporal attentional window in natural vision. The controls (comparisons of neuronal response during looks to targets vs distractors and looks towards and away from the target) and statistical rigor make these findings compelling. There will likely be additional future impacts of this study. For example, the eye movement patterns collected in this study may also provide a valuable dataset for future study of understanding search strategies. Goal-directed vs non-goal-directed task comparisons could be designed to test possible circuit models. Although much remains unknown regarding how and where frontal and temporal signals are integrated during active search, these data contribute important guideposts for future models of active visual search.

    3. Reviewer #3 (Public review):

      In this manuscript, the authors investigate the role of attention in foveal processing during a naturalistic task. They record neural activity from extrastriate visual areas V4 and inferotemporal cortex, as well as from the lateral prefrontal cortex, in macaques performing a free-gaze visual search task. In this task, animals searched for a face or house target among multiple complex stimuli, with no constraints on eye movements. Unlike classic studies of visual attention, which often rely on controlled fixation, this work examines neural activity in both foveal and peripheral receptive fields during naturalistic eye movements.

      The main question addressed by the authors is how feature-based attention is distributed and coordinated across foveal and peripheral visual fields during active search, and how this attentional processing influences saccade behavior. The authors show that foveal units in visual areas exhibit feature-based attentional enhancement, with stronger responses when a fixated stimulus is a target compared to when the same stimulus serves as a distractor. Peripheral units in visual and prefrontal areas show both feature-based and spatial attentional modulation, consistent with prior work. Finally, the authors show that attentional modulation depends primarily on stimulus category rather than response magnitude, with neurons showing similar enhancement for all images within the target category regardless of how strongly individual images drive the cell.

      There are several notable strengths of this paper including:

      (1) Disentangling feature-based and spatial attention during naturalistic vision remains a central challenge. This paper tackles both simultaneously, parsing neural populations by object selectivity (face-selective, house-selective, non-selective) and RF position (foveal vs. peripheral).

      (2) The unconstrained search task (Fig. 1A) moves beyond the dominant fixed-gaze, cued-attention designs (Zhou & Desimone, 2011) to study attention as it operates during natural behavior, with sequential fixations and voluntary saccades.

      (3) The scale of the multi-area recordings is a major strength and is well aligned with current trends in primate and human neuroscience toward large-scale, multi-area recordings. Simultaneous recordings from visual and prefrontal areas, comprising over 4,900 foveal units and more than 1,500 peripheral units, enable meaningful cross-area latency comparisons and area-specific analyses of attentional modulation. This study builds on the authors' previous analyses of this dataset by expanding the scope to show that feature-based attention generalizes across neuronal classes and operates on categorical identity rather than response magnitude.

      (4) The combination of simultaneous multi-area recordings and a rich behavioral paradigm provide a dataset that is well suited for population decoding, cross-area interaction analyses, and trial-by-trial prediction of saccade choices, which could substantially deepen mechanistic understanding beyond the largely univariate comparisons presented here.

      While the data broadly support the paper's main conclusions, several issues limit the strength of the mechanistic interpretation and should be taken into consideration:

      (1) Receptive field size is not explicitly quantified and may confound foveal-peripheral comparisons. Units are classified as foveal or peripheral based on responsiveness to the cue versus the search array (Methods, p. 17), but the manuscript lacks essential information about receptive field sizes, eccentricities, and the number of search stimuli falling within each receptive field and related proper controls. This is critical because receptive fields in visual area V4 at foveal eccentricities are relatively small (Gattass et al., 1988; Desimone & Schein, 1987), whereas receptive fields in inferotemporal cortex can span several degrees to tens of degrees and often include the fovea (Op de Beeck & Vogels, 2000; DiCarlo & Maunsell, 2003; Zoccolan et al., 2007). Given the 2{degree sign} × 2{degree sign} stimulus size, multiple search items could potentially fall simultaneously within peripheral receptive fields. This introduces a potential confound, as attentional modulation is known to be strongest when multiple stimuli appear within a single receptive field (Reynolds et al., 1999). Although the authors acknowledge this issue for visual area V4 (p. 17), it is neither quantified nor controlled for. Without explicit receptive field mapping relative to the search array, comparisons between foveal and peripheral units, as well as between visual areas, are difficult to interpret cleanly.

      (2) Attentional modulation is difficult to dissociate from saccade planning and decision-related signals. The free-gaze paradigm enhances ecological validity but introduces a temporal confound: mean distractor fixation durations are approximately 156 ms (p. 9), while attentional effects emerge between 137 and 170 ms after fixation onset (Fig. 2). As a result, the reported attentional modulation coincides with preparation of the subsequent saccade. Neural activity measured in the primary analysis window (150-225 ms; p. 19) therefore likely reflects a mixture of visual, attentional, motor planning, target recognition, and behavioral relevance signals, all of which are known to modulate responses in visual areas at similar latencies (e.g., Chelazzi et al., 1998). Moreover, target fixations (~257 ms) and distractor fixations (~156 ms) occur on fundamentally different behavioral timescales, which may inflate apparent foveal attentional effects. While the authors suggest that these timing differences support the idea that foveal feature-based attention facilitates prolonged fixation on target stimuli, this interpretation is not fully supported by the current analyses. That said, the saccade-aligned analyses of peripheral units (Fig. S3) partially mitigate this concern by demonstrating that feature-based modulation persists through saccade execution.

      (3) The "attention-out" condition for spatial attention lacks directional control. In the spatial attention analyses (Fig. 4D-F), the "attention-out" condition appears to include all fixations followed by saccades directed away from the receptive field, regardless of saccade direction. This differs from classic spatial attention designs, which typically use controlled anti-saccades or saccades to fixed locations opposite the receptive field (e.g., Moore & Armstrong, 2003; Gregoriou et al., 2009). Saccades directed toward locations adjacent to, but outside, the receptive field may still partially engage spatial attention mechanisms near the receptive field via broad attentional fields or motor preparation gradients (Bisley & Goldberg, 2010). In addition, the "attention-out" condition likely contains a heterogeneous mixture of trials in which the stimulus in the receptive field is either a target or a distractor, since feature-based attention effects are derived from this same pool of trials. As a result, spatial and feature attention effects are not fully orthogonal, and variance related to feature attention may already be embedded in the spatial attention baseline.

      [Editors' note: the authors have provided responses to each of these points.]

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript aims to differentiate between foveal and peripheral attentional mechanisms in visual and frontal brain regions in monkeys engaged in a free-gaze visual search task.

      Strengths:

      The manuscript is clearly written, the question is important, and the behavioral task is interesting.

      Weaknesses:

      I have two major concerns.

      (1) The authors interpret divergence in neural responses to target vs nontarget as attention. But it is not. The subject has to attend to both target and nontarget stimuli to determine the stimulus category and thereby decide on the next action. Thus, divergence between target and nontarget responses could reflect categorical discrimination, but I am not sure this can be interpreted as attentional modulation. While it may be tempting to suggest that finding a stimulus of a specific category is "feature attention", analogous to, e.g., attending to the red stimulus, I don't believe this is correct. For the former, the animals have to attend to a stimulus, and examine the stimulus to determine the stimulus category, unlike a simpler discrimination, which may pop out. Given this, I am unconvinced that the interpretations in this manuscript are valid.

      We thank the reviewer for raising this concern. Selective attention is a process of focusing on goal-relevant stimuli (targets) while ignoring irrelevant distractions. Importantly, attentional selection is not limited to simple visual features (e.g., color, shape, or motion); it can also operate over more complex features. For example, objects themselves can serve as units of attentional selection [1, 2], and feature-based attentional effects have been observed when searching for images that match the cued images or image patches [3, 4]. In this context, attention can be directed either to overall features of an object or to objects as configurations of multiple non-spatial features. Furthermore, attention to the category of stimuli has been extensively investigated in fMRI experiments in humans [5-8], and it has been shown that attention can warp the representations of semantically related categories when participants search for different categories [7].

      Similarly, in our study, monkeys were trained to search for images that matched the category of the cue. The neural responses to targets versus distractors were compared while constrained to the same stimuli across different trials, ensuring that the observed response divergence was not due to the physical category of the targets and distractors. We also included only neural responses occurring prior to fixations associated with target selection, that is, before the monkeys made a behavioral choice, thereby controlling for potential contributions of target detection or decision-related signals to the observed effects.

      We have clarified and addressed this point in the Discussion as follows:

      “Feature-based attention to simple visual features such as color, shape, or motion has been extensively studied [1, 3, 5, 7-9, 11, 12, 64]. Attention can also operate over more complex features. For example, objects themselves can serve as units of attentional selection [65, 66], and feature-based attentional effects have been observed when searching for images that match the cued images or image patches [6, 67]. In this context, attention can be directed either to overall features of an object or to objects as configurations of multiple non-spatial features. Furthermore, attention to the category of stimuli has been extensively investigated in fMRI experiments in humans [68-71], and it has been shown that attention can warp the representations of semantically related categories when participants search for different categories [70]. In this study, the neural responses to targets versus distractors were compared while constrained to the same stimuli across different trials, ensuring that the observed response divergence was not due to the physical category of the targets and distractors.”

      (2) Regarding the RF classification of foveal and peripheral RFs for IT and PFC, prior work suggests that neurons in IT cortex (especially AIT) and PFC have RFs that largely include the foveal visual field. So, it would be important to include figures that show the RFs of neurons classified as foveal versus peripheral for all three areas.

      We thank the reviewer for raising this important point. We agree with the reviewer that neurons in IT cortex and PFC often have RFs that include the foveal visual field. We did record foveal units with both focal and broad foveal RFs; however, in our analysis we only included neurons with focal foveal RFs to exclude the influence of peripheral stimuli. We defined focal foveal-RF units as those that responded solely to the cue in the foveal region and not to items in the search array presented at least 5° away from the central fixation point, ensuring that their RFs did not extend to these peripheral locations. The items were also separated by at least 5° from each other, excluding the possibility that peripheral stimuli fell within their RFs during fixations. By definition, their RFs were restricted to the central point. This is further supported by Fig. S1A-H, which shows no responses to items in the search array at peripheral locations. We have made modifications in the Results and Methods as follows:

      “Notably, the items in the search array were presented at least 5° from the central fixation point and were also separated by at least 5° from each other, excluding the possibility that peripheral stimuli fell within their foveal RFs during fixations.”

      And:

      “In this study, our focus was on units with focal foveal RFs and units with localized peripheral RFs. All further analyses were conducted on these units.”

      We modified Fig. 1 to illustrate the RFs of neurons classified as peripheral, which were also characterized in our previous study using the same dataset [9]. The peripheral population exhibits no responses to the central cue (Fig. S1I–T).

      Reviewer #2 (Public review):

      Summary:

      In natural visual behavior, such as when one is looking for a face in the crowd, the eyes are moved from site to site, seeking possible matching targets. This involves attention both to the current view at the center of vision (the foveal location) as well as to upcoming views via attention to targets in the periphery. While it has been established that attention generally enhances neuronal response (compared to simple visual activation) at the attended spatial location, this study provides solid evidence that attention during active visual search leads to neuronal response enhancement only when the eye moves towards targets that exhibit the desired feature and category. This study thus moves the field towards understanding the neural encoding of active vision.

      This study examines the neuronal basis of feature-selective attention during active, freely behaving visual search. Traditional electrophysiological studies on visual attention in monkeys commonly used an eye fixation with a covert attention paradigm, but have not sufficiently addressed the roles of both foveal and peripheral attention in play during natural looking behavior. Here, the authors present a novel paradigm in which, during eye-movement mediated search, neuronal receptive fields are recorded in multiple cortical areas (sensory V4, temporal, and prefrontal areas). In this manner, as the eye foveates, items in the array fall into foveal or non-foveal recorded sites. Thus, the experimental paradigm is elegant, offering the opportunity to make multiple types of comparisons: target/distractor, towards/away from fovea, and areal. Specifically, following a category cue (face, house, hand, flower), freely initiated saccades are made to locate a categorically matching 'target' in an array of distractors. Feature attention is assessed by comparing eye saccades made to targets vs to distractors. Spatial attention is assessed by comparing saccades made 'towards' vs 'away' from targets. Statistics are rigorous and nicely designed. The detailed association of simultaneously obtained eye movement sequences and neural parameters is well done. These are valuable data that will contribute to our understanding of attentional modulation in visual search.

      Strengths:

      The significance of these findings is fundamental. Decades of attention research in vision have been based on the paradigm of visual fixation and covert peripheral attention. However, increasingly, the field has moved towards understanding how the visual system works during active vision. Here, the authors use an active visual search paradigm and record from multiple areas (V4, IT, PFC). They find enhancement of attention both in the foveal and peripheral locations, and, furthermore, a high degree of feature and categorical specificity. This provides valuable data for the concept of a foveal-peripheral attentional window in natural vision. The controls (comparisons of neuronal response during looks to targets vs distractors, and looks towards and away from the target) and statistical rigor make these findings quite compelling.

      Weaknesses:

      While the study is generally quite strong, there are a few weaknesses to be addressed.

      (1) Little rationale is provided for recording in the selected areas, V4, IT, and PFC. Given the respective roles in sensory, object recognition, and goal-directed behavior, some rationale for this design should be offered, and commonalities/distinctions between these areas should be discussed.

      We thank the reviewer for the suggestion and we modified and added the rationale to the Introduction as follows:

      “V4 and inferotemporal cortex (IT), as the middle and high-level areas of the ventral visual stream, are important for object recognition and categorization [27-34], and their roles have been extensively studied in central vision. At the neuronal level, however, most investigations have largely neglected their functions during active, free-gaze visual search. The prefrontal cortex, including LPFC, has long been implicated as a source of top-down signals that bias the selection of attended features and modulate visual cortical responses [6, 9, 11, 35-40]. Although target-related visual responses have been reported in IT during visual exploration [41], and target-selective responses have been observed in the human medial temporal lobe (MTL) [42] and medial frontal cortex (MFC) [43] during visual search, these studies did not map the receptive fields (RFs) of recorded neurons.”

      We also added a discussion as follows:

      “Some studies have provided evidence for integration between peripheral and foveal feature information across saccades, including features such as stimulus color [58, 59] and object orientation [60, 61], and visual features have been shown to be predictively remapped prior to saccades [62]. Our finding provides a potential neuronal mechanism that may support this integration process [63]. We found that LPFC’s extensive representation of the visual periphery provides a neural substrate for monitoring the broader search array. Crucially, our finding that LPFC activity temporally precedes attentional effects in the visual area—consistent with previous studies [6, 9, 11, 35-40] suggests that it does not merely reflect peripheral sensory input. Instead, LPFC likely acts as a top-down orchestrator, projecting task-relevant templates derived from current foveal goals onto peripheral candidate locations, a possibility that warrants further investigation.”

      (2) Given the reliance of all analyses on saccadic behavior (towards target/distractor, towards/away from target), additional description and summaries of eye movement behavior during single trials and across trials should be provided.

      We thank the reviewer for this helpful suggestion. We have added a description of saccade behavior to the Results as follows:

      “The mean number of saccades monkeys made to find the target after the onset of the search array was 2.25 ± 1.35 (mean ± SD across trials; Table 1) of correct trials, and the mean saccade amplitude was 7.99° ± 3.58° (mean ± SD across saccades; Table 1). Monkeys could fixate on each distractor or the target freely, provided they did not maintain fixation on the target for longer than 800 ms. Across sessions, 42.44% ± 3.6% of saccades were directed to distractors, 57.56% ± 3.6% to targets, and 12.59% ± 3.46% were saccades away from targets (see our previous studies [44-46] for detailed behavioral analyses).”

      We have modified Fig. 1A and its legend to illustrate the saccadic patterns of monkeys during the search task.

      We have also included Table 1, which summarizes eye movement behaviors.

      (3) The dependency of findings on top-down (categorical & feature-specific) task design should be discussed.

      We thank the reviewer for the suggestion and added a discussion as follows:

      “In this task, attention is strongly guided by top-down goals, which bias processing toward behaviorally relevant features and object categories [2, 50, 51]. Top-down attention, including categorical and feature-specific components, has been shown to modulate neural processing across the visual pathway based on task demands and to originate from distributed frontoparietal control networks [11, 35-38, 40]. Our study provides further insight into the mechanisms of goal-directed visual attention, as it is among the first to demonstrate foveal feature attention effects during free-gaze visual search, as well as the distribution of feature and spatial attention across the entire visual field.”

      Reviewer #3 (Public review):

      In this manuscript, the authors investigate the role of attention in foveal processing during a naturalistic task. They record neural activity from extrastriate visual areas V4 and inferotemporal cortex, as well as from the lateral prefrontal cortex, in macaques performing a free-gaze visual search task. In this task, animals searched for a face or house target among multiple complex stimuli, with no constraints on eye movements. Unlike classic studies of visual attention, which often rely on controlled fixation, this work examines neural activity in both foveal and peripheral receptive fields during naturalistic eye movements.

      The main question addressed by the authors is how feature-based attention is distributed and coordinated across foveal and peripheral visual fields during active search, and how this attentional processing influences saccade behavior. The authors show that foveal units in visual areas exhibit feature-based attentional enhancement, with stronger responses when a fixated stimulus is a target compared to when the same stimulus serves as a distractor. Peripheral units in visual and prefrontal areas show both feature-based and spatial attentional modulation, consistent with prior work. Finally, the authors show that attentional modulation depends primarily on stimulus category rather than response magnitude, with neurons showing similar enhancement for all images within the target category regardless of how strongly individual images drive the cell.

      There are several notable strengths of this paper, including:

      (1) Disentangling feature-based and spatial attention during naturalistic vision remains a central challenge. This paper tackles both simultaneously, parsing neural populations by object selectivity (face-selective, house-selective, non-selective) and RF position (foveal vs. peripheral).

      (2) The unconstrained search task (Figure 1A) moves beyond the dominant fixed-gaze, cued-attention designs (Zhou & Desimone, 2011) to study attention as it operates during natural behavior, with sequential fixations and voluntary saccades.

      (3) The scale of the multi-area recordings is a major strength and is well aligned with current trends in primate and human neuroscience toward large-scale, multi-area recordings. Simultaneous recordings from visual and prefrontal areas, comprising over 4,900 foveal units and more than 1,500 peripheral units, enable meaningful cross-area latency comparisons and area-specific analyses of attentional modulation. This study builds on the authors' previous analyses of this dataset by expanding the scope to show that feature-based attention generalizes across neuronal classes and operates on categorical identity rather than response magnitude.

      (4) The combination of simultaneous multi-area recordings and a rich behavioral paradigm provides a dataset that is well-suited for population decoding, cross-area interaction analyses, and trial-by-trial prediction of saccade choices, which could substantially deepen mechanistic understanding beyond the largely univariate comparisons presented here.

      While the data broadly support the paper's main conclusions, several issues limit the strength of the mechanistic interpretation and should be taken into consideration:

      (1) Receptive field size is not explicitly quantified and may confound foveal-peripheral comparisons. Units are classified as foveal or peripheral based on responsiveness to the cue versus the search array (Methods, p. 17), but the manuscript lacks essential information about receptive field sizes, eccentricities, and the number of search stimuli falling within each receptive field and related proper controls. This is critical because receptive fields in visual area V4 at foveal eccentricities are relatively small (Gattass et al., 1988; Desimone & Schein, 1987), whereas receptive fields in inferotemporal cortex can span several degrees to tens of degrees and often include the fovea (Op de Beeck & Vogels, 2000; DiCarlo & Maunsell, 2003; Zoccolan et al., 2007). Given the 2{degree sign} × 2{degree sign} stimulus size, multiple search items could potentially fall simultaneously within peripheral receptive fields. This introduces a potential confound, as attentional modulation is known to be strongest when multiple stimuli appear within a single receptive field (Reynolds et al., 1999). Although the authors acknowledge this issue for visual area V4 (p. 17), it is neither quantified nor controlled for. Without explicit receptive field mapping relative to the search array, comparisons between foveal and peripheral units, as well as between visual areas, are difficult to interpret cleanly.

      We thank the reviewer for the helpful suggestion and apologize for not explicitly providing essential information about the RFs of the units. We added a detailed description of RF properties to the Results as follows:

      “The RFs of these peripheral units were further mapped using a visually guided saccade task and quantified by the number of stimuli that activated each unit (Fig. 1F-K). The eccentricities of the peripheral RFs were 6.22° ± 1.31° (mean ± SD) in V4, 7.04° ± 1.52° in IT, and 6.68° ± 1.56° in LPFC. The sizes of the peripheral RFs were 3.67° ± 1.87° in V4, 6.86° ± 3.11° in IT, and 8.65° ± 3.02° in LPFC. The numbers of items from the search array falling within peripheral RFs were 1.49 ± 0.55 in V4, 2.2 ± 0.72 in IT, and 2.56 ± 0.74 in LPFC (also see our previous study [44]).”

      The reviewer is correct that multiple items from the search array did fall within the RFs of peripheral-RF units. However, for focal foveal units, only the fixated stimulus fell within the RF, due to the design of the search array and the definition of these units used in our analyses (see our reply to Reviewer 1, Public Review, Question 2 for details). We agree with the reviewer that attentional modulation is typically stronger when multiple stimuli fall within RFs. In our design, peripheral RFs, on average, contained more stimuli than foveal RFs. Therefore, this difference in RF size would, if anything, be expected to bias toward stronger attentional modulation in peripheral units. This would make our observation conservative, thereby further supporting rather than undermines our main finding of robust feature-based attentional enhancement in foveal units, challenging the prevailing view that such modulation is predominantly peripheral. However, we agree that, when comparing the latency of attentional effects across brain regions in Fig. 3, we cannot rule out the influence of the number of stimuli arising from differences in RF size.

      (2) Attentional modulation is difficult to dissociate from saccade planning and decision-related signals. The free-gaze paradigm enhances ecological validity but introduces a temporal confound: mean distractor fixation durations are approximately 156 ms (p. 9), while attentional effects emerge between 137 and 170 ms after fixation onset (Figure 2). As a result, the reported attentional modulation coincides with the preparation of the subsequent saccade. Neural activity measured in the primary analysis window (150-225 ms; p. 19), therefore, likely reflects a mixture of visual, attentional, motor planning, target recognition, and behavioral relevance signals, all of which are known to modulate responses in visual areas at similar latencies (e.g., Chelazzi et al., 1998). Moreover, target fixations (~257 ms) and distractor fixations (~156 ms) occur on fundamentally different behavioral timescales, which may inflate apparent foveal attentional effects. While the authors suggest that these timing differences support the idea that foveal feature-based attention facilitates prolonged fixation on target stimuli, this interpretation is not fully supported by the current analyses. That said, the saccade-aligned analyses of peripheral units (Figure S3) partially mitigate this concern by demonstrating that featurebased modulation persists through saccade execution.

      We thank the reviewer for raising this important question. We agree that the temporal overlap of visual, motor planning, target recognition, and behavioral relevance signals with attention can result in mixed activity, which needs to be dissociated. Therefore, when calculating feature-based attention, we did implement a series of controls. We added a discussion as follows:

      “A major challenge in interpreting neural activity related to attentional modulation is the inherent temporal overlap of visual processing, motor planning, and target recognition signals in the free-gaze visual search task [73]. To isolate genuine feature-based attention from potential confounds, we applied several stringent analytical constraints, consistent with prior studies [3, 5, 6]. Specifically, by restricting our analysis to fixations where the subsequent saccade was directed away from the RFs, we dissociated attentional modulation from the preparatory motor activity associated with saccade execution. Furthermore, by comparing responses to the same physical stimulus alternating its role as a target or distractor across trials we eliminated any potential bias introduced by stimulus identity or physical category. We restricted our analysis to fixations preceding target selection that is, before the monkeys made a behavioral choice to minimize contributions from target detection or decision-related signals.”

      We thank the reviewer for pointing out the issue of different timescales for target versus distractor fixations. To address this, we conducted a control analysis by computing foveal feature-based attentional modulation using fixations on targets and distractors with matched fixation durations. We obtained similar results. We have updated Fig. S2 to include this control analysis.

      We also clarified this point in the Results as follows:

      “We also obtained similar results when controlling for fixation durations on targets and distractors (i.e., there was no significant difference between fixation durations on targets and distractors; Wilcoxon signed-rank test, P > 0.05; Fig. S2K–P).”

      Lastly, as the reviewer correctly pointed out, the interpretation that foveal feature-based attention facilitates prolonged fixation on the target was not supported. We have revised the Results as follows:

      “On average, target fixations (256.69 ± 197.44 ms [mean ± SD]) were significantly longer than distractor fixations (156.26 ± 45.94 ms; Wilcoxon rank-sum test, P < 0.0001), and during these prolonged target fixation, foveal feature-based attention modulation was consistently observed.”

      (3) The "attention-out" condition for spatial attention lacks directional control. In the spatial attention analyses (Figures 4D-F), the "attention-out" condition appears to include all fixations followed by saccades directed away from the receptive field, regardless of saccade direction. This differs from classic spatial attention designs, which typically use controlled anti-saccades or saccades to fixed locations opposite the receptive field (e.g., Moore & Armstrong, 2003; Gregoriou et al., 2009). Saccades directed toward locations adjacent to, but outside, the receptive field may still partially engage spatial attention mechanisms near the receptive field via broad attentional fields or motor preparation gradients (Bisley & Goldberg, 2010). In addition, the "attention-out" condition likely contains a heterogeneous mixture of trials in which the stimulus in the receptive field is either a target or a distractor, since feature-based attention effects are derived from this same pool of trials. As a result, spatial and feature attention effects are not fully orthogonal, and variance related to feature attention may already be embedded in the spatial attention baseline.

      We thank the reviewer for this important question. We performed a directional control analysis by computing spatial attentional modulation using paired fixations from the attention-in and attention-out conditions. Only saccades directed in nearly opposite directions—defined as having a saccade direction angle ≥ 170° within the 0–180° range—were included. We obtained similar results (Author response image 1). 

      Author response image 1.

      Peripheral spatial attentional modulation in V4, IT, and LPFC. Population response to stimuli followed by saccades directed into their RFs (attention in) versus directed approximately opposite and outside their RFs (attention out), shown for V4 (A), IT (B), and LPFC (C). Shaded area denotes ±SEM across units.

      We did control for feature-based attention when calculating spatial attentional modulation. We apologize for the lack of clarity and have added a description of this control to the Methods as follows:

      “The saccade-target stimulus in the RF during attention-in fixations was matched to a stimulus in the same location during attention-out fixations; in both conditions, this stimulus always served as a distractor for that trial, except in the “Distractor fixations to T” condition (Fig. 5 and Fig. S4), in which it instead served as the target. This design eliminates differences due to feature-based attention between the attention-in and attention-out conditions.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Figure 3C: Unclear how to compare LPFC vs V4 for foveal units since only data from peripheral LPFC is shown?

      We thank the reviewer for pointing out this mistake. In Fig. 3C, we only compared LPFC peripheral units, V4 peripheral units, and V4 foveal units. We have corrected this in the legend of Fig. 3 as follows:

      “Shown are cumulative distributions of feature-attention effect latencies, computed from individual foveal face-, house-, and non-selective units in V4 and IT, and from peripheral non-selective units in V4,

      IT, and LPFC.”

      (2) On page 8, last para: For units with peripheral RFs ... Is this controlled for whether the saccade is to targets or to distractors?

      We thank the reviewer for the question. We indeed addressed this concern by separating fixations based on whether the subsequent saccade was directed to a target or a distractor, and by analyzing attention modulation within each condition. Therefore, attention effects were evaluated while holding the saccade destination constant, effectively controlling for potential confounds related to saccade target selection.

      (3) Page 9: The authors find that target fixations were longer than distractor fixations and conclude that this supports the idea that foveal feature-based attention increases fixation duration, but this interpretation is pure conjecture, and there is no experimental manipulation presented in this paper that helps to establish this interpretation.

      We thank the reviewer for this important comment. We agree that this observation does not, by itself, support our original interpretation, and we have modified it in the Results. Please refer to the last paragraph of our Reply to Question 2 from Reviewer 3 (Public Review).

      (4) Data analysis: receptive field. The authors state that visual response to a cue and the stimulus array was assessed during the 0-200 ms window after stimulus onset. However, after the array onset, the animal could saccade within the 200 ms window. How do the authors ensure uniform stimulation during the 0-200 ms window?

      We thank the reviewer for this question. The activity of units in V4, IT, and LPFC within the 200 ms window after array onset primarily reflected visual stimulation prior to saccades, because typical saccade latencies were approximately 150–200 ms, and the response onset latencies of these units were around 50 ms.

      (5) On page 19, the authors state that to assess feature attention in peripheral RFs, they divided trials into target and distractor fixations. In the former, there was a target in the neuron's RF. This is confusing. I assume target fixations imply fixating on a target, but the authors may mean fixations where a target is in the RF. Please clarify.

      We thank the reviewer for pointing out this confusion. In the original manuscript, we intended to sort fixations by whether a target stimulus was located within the unit’s peripheral RF. To avoid further confusion, we have revised the description in the Methods as follows:

      “we sorted fixations during the search period, following a procedure similar to that in our previous study [5], into two types: “target” – a target stimulus was located within the unit’s peripheral RF; and “distractor” – the same stimulus appeared in the same peripheral RF location but served as a distractor.”

      (6) Figure S1: Are these example units? How many trials? SEM? The sharp rise and no noise are inconsistent; the former suggests minimal smoothing, while the latter suggests lots of smoothing.

      We thank the reviewer for these questions. We showed average responses across all units in Fig. S1. On average, there were 941.79 ± 182.56 trials (mean ± SD across sessions). Shaded areas indicate ±SEM across units. The sharp rise reflects the synchronous response of neurons to the stimulus, while the smooth appearance and low noise result from averaging across a very large number of units and trials.

      Reviewer #2 (Recommendations for the authors):

      Major comments:

      (1) One weakness of this manuscript is the lack of a rationale for choosing V4, IT, and PFC. Specifically, what are the predictions of the roles of these respective areas in the integration of current and peripheral (future foveal) views? There is a significant literature linking the pre-saccadic peripheral stimulus and the post-saccadic foveal stimulus, suggesting that both spatial and temporal integration occur. However, whether such integration occurs at high or low cortical levels is unknown. By recording from mid-tier (V4) and high-order areas (IT, PFC), the authors have an opportunity to address this question. However, there is no mention of this topic, either in the introduction, results, or discussion. I find this omission surprising. At the very least, it should contribute to experimental design rationale and some discussion.

      We thank the reviewer for the suggestion and we modified and added the rationale to the Introduction and a discussion about this integration. Please refer to our Reply to Question 1 from Reviewer 2 (Public Review).

      (2) As both behavior and neural recordings are collected, a figure on saccadic patterns would enhance the reader's understanding. Questions that come to mind are: What does a single search trial look like? How many saccades are there per trial? How often is the target identified after 1, 2, 3, etc saccades? What is the average size of a saccade? Although this is not a study of search strategy per se, a modicum of description of the search sequences would provide context on the behavior. I suggest an illustration of one or more sample trials; a summary of saccade behavior would also be helpful for understanding the data in relation to behavioral performance.

      We thank the reviewer for this helpful suggestion. We have modified Fig. 1A and its legend to illustrate the saccadic patterns of monkeys during the search, providing an example of a single search trial. Additionally, we have added a description of saccade behavior to the Results and included Table 1, which summarizes eye movement behavior. Please refer to our Reply to Question 2 from Reviewer 2 (Public Review) for further details.

      (3) "Consistently, the probability of making a saccade to a peripheral target was higher following distractor fixations (75.22%) than following target fixations (48.44%, or 63.49% after probability calibration; see Methods), indicating the important role of peripheral feature-based attention in guiding eye movements" It should be noted that this target-oriented visual search is fundamentally a top down task. Once the target is found, the reward is obtained; saccades to distractors are not rewarded, so saccades are more likely. So certainly this task design would increase the post-distractor saccades and decrease the number of post-target saccades. Please clarify the behavioral paradigm: once a reward is obtained, does the task continue, or is a new trial initiated?

      We apologize for the confusion regarding the behavioral paradigm. We would like to clarify that when the target was found and fixated for 800 ms, the reward was delivered and no further saccades occurred. However, if the target was not fixated for 800 ms, the search could continue. It is worth noting that the target fixations in our analyses were restricted to those occurring during ongoing search behavior, excluding target fixations associated with trial termination and reward delivery. Moreover, we compared the probability of making a saccade to the target, rather than the absolute number of saccades, following these fixations. We have modified the Results for clarification, as follows:

      “Two monkeys performed a category-based visual search task, where their objective was to fixate on one of the two search targets that matched the category of the cue (Fig. 1A, B). Specifically, the monkeys were presented with a central fixation point for 400 ms, followed by a cue lasting 500-1300 ms. After a 500 ms delay, a search array appeared with 11 items, including two targets, randomly chosen from 20 possible locations (Fig. 1E). The monkeys had 4000 ms to find one target and maintain fixation on it for 800 ms to earn a juice reward. Fixating on either target completed the trial, and the monkeys did not search for the second target. A new trial began after the reward. It is worth noting that the two target stimuli matched the category of the cue but were different images. The monkeys were required to maintain fixation throughout the cue and delay periods. During search, however, eye movements were unconstrained, and monkeys could revisit each search distractor or target as long as they did not fixate on a target for 800 ms.”

      (4) The fact that there are many more peripheral units in LPFC suggests that this is a region of foveal/periph integration. Combined with the finding that the LPFC leads the attentional effects, this should be a discussion point.

      We thank the reviewer for the suggestion and we added a discussion as follows:

      “Some studies have provided evidence for integration between peripheral and foveal feature information across saccades, including features such as stimulus color [58, 59] and object orientation [60, 61], and visual features have been shown to be predictively remapped prior to saccades [62]. Our finding provides a potential neuronal mechanism that may support this integration process [63]. We found that LPFC’s extensive representation of the visual periphery provides a neural substrate for monitoring the broader search array. Crucially, our finding that LPFC activity temporally precedes attentional effects in the visual area consistent with previous studies [6, 9, 11, 35-40] suggests that it does not merely reflect peripheral sensory input. Instead, LPFC likely acts as a top-down orchestrator, projecting task-relevant templates derived from current foveal goals onto peripheral candidate locations, a possibility that warrants further investigation.”

      Minor comments:

      (1) Figures 2A-D. "These face-selective units also showed slightly enhanced responses to house targets in IT (P < 0.05), but not in V4 (P = 0.89)." It does not appear enhanced.

      We agree with the reviewer that the effect is modest and does not appear strongly enhanced. However, the average response in the 150–225 ms time window to the house target was significantly higher than that to the house distractor in IT face-selective units (Wilcoxon signed-rank test, P = 0.042). We modified the description in the Results as follows: 

      “These face-selective units also showed weakly but significantly enhanced responses to house targets in IT (P < 0.05)”

      (2) Figure 3. For population comparison, a bootstrapped null distribution was used, and a 2-sided permutation test was used to determine the latency difference between the target and distractor; please show these results (described in text) in a figure. Figures 3A-C are described as the latency of individual units. So each of these graphs is the mean of multiple units? So this is also a population analysis? What is the difference between these two comparisons? This is somewhat confusing.

      We apologize for the confusion and thank the reviewer for pointing this out. Each panel in Fig. 3 shows the cumulative distribution of latencies across individual units within each brain region, reflecting the variability of response timing across single neurons. For this analysis, we first calculate the latency of each unit separately. In contrast, population-level latency is measured from the averaged responses of all units within each region (Fig. 2), which captures the overall timing of the population response rather than individual variability. Statistical comparisons at the population level are performed using a two-sided permutation test. We modified Fig. 2 to better illustrate the population-level latency results.

      (3) Did peripheral RFs span more than a single stimulus in the array? If so, how does this impact the interpretation of Figure 5?

      We thank the reviewer for pointing this out. The reviewer is correct that, in peripheral RFs, more than one stimulus from the search array could fall within the receptive field (1.49 ± 0.55 in V4, 2.2 ± 0.72 in IT, and 2.56 ± 0.74 in LPFC). We controlled for this in our analysis of both feature-based and spatial attention effects for peripheral units in Fig. 5. For feature-based attention, we performed the analysis in a stimulus-by-stimulus manner within each category (house and face), such that when a given stimulus served as the target, it was the only target within the RF, and when it served as a distractor, it was the only distractor of its category within the RF. Although additional distractor could still fall within the RF, their identities were random across conditions and thus would be averaged out. A similar approach was applied to spatial attention, where the stimulus-by-stimulus comparison was extended across all four categories, and attention-out stimuli were paired with the corresponding saccade-target stimuli in the attention-in condition, with the effects of other randomly present distractors averaged out. Therefore, the effects shown in Fig. 5 reflect comparisons at the level of individual stimulus, minimizing confounds from other stimuli within the RF.

      (4) Figure 5G: "during "Target fixations to D", there was no significant feature attentional enhancement in response to the peripheral target (Wilcoxon signed-rank test, P > 0.05; Figure 5G-I left panels). It appears that there is some effect of spatial attention during Target Fix to D trials.

      We thank the reviewer for pointing this out and have revised the Results as follows:

      “We further found that spatial attentional enhancements to the saccade target were reduced during target fixations compared to distractor fixations in V4 and IT when activity was aligned to fixation onset (Wilcoxon rank-sum test, P < 0.05; Fig. 5G, H versus Fig. 5A, B), although this effect was not completely abolished.”

      (5) The specific areas of IT and LPFC that were recorded should, as much as possible, be mentioned.

      We thank the reviewer for the helpful suggestions and have added a description of the specific IT and LPFC recording sites to the Methods as follows:

      “Recordings in IT spanned the central IT cortex, encompassing the area between the anterior middle temporal sulcus (AMTS) and the posterior middle temporal sulcus (PMTS), including TE and TEO. Recordings in LPFC were located anterior to the arcuate sulcus (AS) and lateral to the principal sulcus (PS), mainly covering areas 45 and 44.”

      (6) It is often difficult to distinguish the different lines, e.g., red solid vs red dotted, due to their overlap. Would the removal of the error band make this clearer? If so, could put full figure with error bands in the Supplementary Figure.

      We thank the reviewer for this helpful suggestion. To improve visual clarity, we adjusted Fig. 6, Fig. 7, Fig. S2, Fig. S3, Fig. S4, and Fig. S6 by changing the line styles and placing the shaded error bands beneath the traces, allowing the lines to remain clearly visible despite overlap.

      (7) For easy access, the number of saccades to/from targets/distractors should be put into a table.

      We thank the reviewer for the suggestion. We calculated the probability of saccades to and from targets and distractors for each session and report the mean ± SD across sessions in Table 1, as the mean number of saccades per trial was only 2.3. Please refer to our Reply to Question 2 from Reviewer 2 (Public Review) for Table 1.

      Reference

      (1) O'Craven, K.M., P.E. Downing, and N. Kanwisher, fMRI evidence for objects as the units of attentional selection. Nature, 1999. 401(6753): p. 584-7.

      (2) Baldauf, D. and R. Desimone, Neural mechanisms of object-based attention. Science, 2014. 344(6182): p. 424-7.

      (3) Hayden, B.Y. and J.L. Gallant, Combined effects of spatial and feature-based attention on responses of V4 neurons. Vision Res, 2009. 49(10): p. 1182-7.

      (4) Bichot, N.P., et al., A Source for Feature-Based Attention in the Prefrontal Cortex. Neuron, 2015. 88(4): p. 832-844.

      (5) Reddy, L. and N. Kanwisher, Category selectivity in the ventral visual pathway confers robustness to clutter and diverted attention. Curr Biol, 2007. 17(23): p. 2067-72.

      (6) Peelen, M.V., L. Fei-Fei, and S. Kastner, Neural mechanisms of rapid natural scene categorization in human visual cortex. Nature, 2009. 460(7251): p. 94-7.

      (7) Cukur, T., et al., Attention during natural vision warps semantic representation across the human brain. Nat Neurosci, 2013. 16(6): p. 763-70.

      (8) Keller, A.S., et al., Attention enhances category representations across the brain with strengthened residual correlations to ventral temporal cortex. Neuroimage, 2022. 249: p. 118900.

      (9) Zhang, J., et al., Behavioral and Neural Mechanisms of Face-Specific Attention during GoalDirected Visual Search. The Journal of Neuroscience, 2024. 44(46): p. e1299242024.

      (10) Bichot, N.P., A.F. Rossi, and R. Desimone, Parallel and serial neural mechanisms for visual search in macaque area V4. Science, 2005. 308(5721): p. 529-534.

      (11) Bichot, N.P., et al., The role of prefrontal cortex in the control of feature attention in area V4. Nat Commun, 2019. 10(1): p. 5727.

      (10) Cohen, M.R. and J.H. Maunsell, Using neuronal populations to study the mechanisms underlying spatial and feature attention. Neuron, 2011. 70(6): p. 1192-204.

      (11) Maunsell, J.H. and S. Treue, Feature-based attention in visual cortex. Trends Neurosci, 2006. 29(6): p. 317-22.

      (12) McAdams, C.J. and J.H. Maunsell, Attention to both space and feature modulates neuronal responses in macaque area V4. J Neurophysiol, 2000. 83(3): p. 1751-5.

      (13) Motter, B.C., Saccadic momentum and attentive control in V4 neurons during visual search. J Vis, 2018. 18(11): p. 16.

      (14) Sapountzis, P., S. Paneri, and G.G. Gregoriou, Distinct roles of prefrontal and parietal areas in the encoding of attentional priority. Proc Natl Acad Sci U S A, 2018. 115(37): p. E8755-E8764.

      (15) Treue, S. and J.C. Martinez Trujillo, Feature-based attention influences motion processing gain in macaque visual cortex. Nature, 1999. 399(6736): p. 575-9.

      (16) Zhou, H. and R. Desimone, Feature-based attention in the frontal eye field and area V4 during visual search. Neuron, 2011. 70(6): p. 1205-17.

    1. eLife Assessment

      One-carbon tetrahydrofolate metabolism plays a crucial role in producing essential metabolic intermediates. In this valuable study, the authors employ a solid genetics-based approach to demonstrate that three distinct metabolic pathways are essential for synthesising 1C-tetrahydrofolates (1C-THF). Disrupting any of these pathways impairs both growth and virulence.

    2. Reviewer #1 (Public review):

      Summary:

      This study identifies three redundant pathways-glycine cleavage system (GCS), serine hydroxymethyltransferase (GlyA), and formate-tetrahydrofolate ligase/FolD-that feed the one-carbon tetrahydrofolate (1C-THF) pool essential for Listeria monocytogenes growth and virulence. Reactivation of the normally inactive fhs gene rescues 1C-THF deficiency, revealing metabolic plasticity and vulnerability for potential antimicrobial targeting.

      Strengths:

      (1) Novel evolutionary insight-Reversible reactivation of a pseudogene (fhs) shows adaptive metabolic plasticity, relevant for pathogen evolution.

      (2) They systematically combine targeted gene deletions with suppressor screening to dissect the folate/one-carbon network (GCS, GlyA, Fhs/FolD).

    3. Reviewer #3 (Public review):

      Summary:

      In this study, Freier et al., demonstrate that 3 distinct metabolic pathways are critical for the synthesis of 1C-THF, a metabolite that is crucial for the growth and virulence of Listeria monocytogenes. Using an elegant suppressor screen, they also demonstrate the hierarchical importance of these metabolic pathways with respect to the biosynthesis of 1C-THF.

      Strengths:

      This study uses elegant bacterial genetics to confirm that 3 distinct metabolic pathways are critical for 1C-THF synthesis in L. monocytogenes and lack of either one of these pathways compromises bacterial growth and virulence. The study uses a combination of in vitro growth assays, macrophage-CFU assays and murine infection models to demonstrate this.

      Comments on revisions:

      The revised manuscript is improved, and the additional genetic experiments provide further support for the proposed metabolic model. However, the central conclusion is not fully established without direct measurement of 1C-THF levels. While I appreciate the authors' explanation regarding the technical limitations, quantitative metabolite measurements (e.g., by mass spectrometry) would have provided much stronger evidence linking the genetic perturbations to altered 1C-THF pools.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study identifies three redundant pathways-glycine cleavage system (GCS), serine hydroxymethyltransferase (GlyA), and formate-tetrahydrofolate ligase/FolD-that feed the one-carbon tetrahydrofolate (1C-THF) pool essential for Listeria monocytogenes growth and virulence. Reactivation of the normally inactive fhs gene rescues 1C-THF deficiency, revealing metabolic plasticity and vulnerability for potential antimicrobial targeting

      Strengths:

      (1) Novel evolutionary insight - reversible reactivation of a pseudogene (fhs) shows adaptive metabolic plasticity, relevant for pathogen evolution.

      (2) They systematically combine targeted gene deletions with suppressor screening to dissect the folate/one-carbon network (GCS, GlyA, Fhs/FolD).

      Weaknesses:

      (1) The study infers 1C-THF depletion mostly genetically and indirectly (growth rescue with adenine) without direct quantification of folate intermediates or fluxes. Biochemical confirmation, LC-MS-based metabolomics of folates/1C donors, or isotopic tracing would strengthen mechanistic claims.

      We agree with the reviewer that quantification of 1C-THF intermediates would strengthen our conclusions. However, the chemical methodologies to extract folates from L. monocytogenes are not established in our lab. Moreover, quantification of C1-substituted folates requires comprehensive biochemical and analytical expertise that we also do not have and which we cannot cover though co-operations. However, to further strengthen our arguments, we have introduced an experiment in the updated manuscript that demonstrates synthetic lethality of a ΔgcvPAB ΔglyA mutant with a deletion of fold (Fig. 7A). This gene encodes 5,10-methylene-tetrahydrofolate dehydrogenase/ 5,10-methylene-tetrahydrofolate cyclohydrolase, which is the third enzyme involved in N5, N10-methylene-THF generation next to GlyA and GcvPBA. Synthetic lethality of a ΔgcvPAB ΔglyA double mutant with a fold deletion is best explained by GcvPAB and GlyA also feeding the N5,N10-methylene-THF pool.

      (2) In multiple result sections, the authors report data from technical triplicates but do not mention independent biological replicates (e.g., Figure 2C, Figure 4A-B, Figure 6D). In addition, some results mention statistical significance but without a detailed description of the specific statistical tests used or replicates, such as Figure 2A-C, Figure 2E, and Figure 2G-I.

      We thank the reviewer for this helpful comment. Experiments were usually repeated three independent times, with each repetition including three technical replicates. Mean values and standard deviations were usually calculated from the technical replicates of a representative run. Statistical significance was calculated using t-tests for pairwise comparisons or t-tests using Bonferroni-Holm correction for multiple comparisons. We made sure that this is explicitly explained for each experiment in the figure legends. Wherever other calculations were used, we also clarified this in the legends.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Freier et al examines the impact of deletion of the glycine cleavage system (GCS) GcvPAB enzyme complex in the facultative intracellular bacterial pathogen Listeria monocytogenes. GcvPAB mediates the oxidative decarboxylation of glycine as a first step in a pathway that leads to the generation of N5, N10-methylene-Tetrahydrofolate (THF) to replenish the 1-carbon THF (1C-THF) pool. 1C-THF species are important for the biosynthesis of purines and pyrimidines as well as for the formation of serine, methionine, and N-formylmethionine, and the authors have previously demonstrated that gcvPAB is important for bacterial replication within macrophages. A significant defect for growth is observed for the gcvPAB deletion mutant in defined media, and this growth defect appears to stem from the sensitivity of the mutant strain to excess glycine, which is hypothesized to further deplete the 1C-THF pool. Selection of suppressor mutations that restored growth of gcvPAB deletion mutants in synthetic media with high glycine yielded mutants that reversed stop codon inactivation of the formatetetrahydrofolate ligase (fhs) gene, supporting the premise that generation of N10-formyl-THF can restore growth. Mutations within the folk, codY, and glyA genes, encoding serine hydroxymethyltransferase, were also identified, although the functional impact of these mutations is somewhat less clear. Overall, the authors report that their work identifies three pathways that feed the 1C-THF pool to support the growth and virulence of L. monocytogenes and that this work represents the first example of the spontaneous reactivation of a L. monocytogenes gene that is inactivated by a premature stop codon.

      Strengths:

      This is an interesting study that takes advantage of a naturally existing fhs mutant Listeria strain to reveal the contributions of different pathways leading to 1C-THF synthesis. The defects observed for the gcvPAB mutant in terms of intracellular growth and virulence are somewhat subtle, indicating that bacteria must be able to access host sources (such as adenine?) to compensate for the loss of purine and fMet synthesis. Overall, the authors do a nice job of assessing the importance of the pathways identified for 1C-THF synthesis.

      Weaknesses:

      (1) Line 114 and Figure 1: The authors indicate that the gcvPAB deletion forms significantly fewer plaques in addition to forming smaller plaques (although this is a bit hard to see in the plaque images). A reduction in the overall number of plaques sounds like a bacterial invasion defect - has this been carefully assessed? The smaller plaque size makes sense with reduced bacterial replication, but I'm not sure I understand the reduction in plaque number.

      The observation that the ΔgcvPAB mutant forms fewer plaques was not our claim, and we have already addressed the possibility of an invasion defect by quantifying bacterial numbers during infection of 3T3 cells. As shown in Fig. 2A, the ΔgcvPAB mutant invades 3T3 cells (the same cells used in the plaque formation assays) as efficiently as the wild type but exhibits reduced intracellular growth. Therefore, the plaquing defect is not due to impaired invasion. Furthermore, we also have analyzed the intracellular dissemination of the ΔgcvPAB mutant in 3T3 fibroblasts compared to a ΔactA mutant by microscopy. This shows that the ΔgcvPAB is evenly distributed throughout the infected host cells as the wild type and unlike the ΔactA mutant, which forms clusters (Fig. S1). Both experiments indicate that the reduced plaque area results from impaired intracellular growth rather than defects in invasion or cell-to-cell spread. The apparent reduction in plaque numbers in ΔgcvPAB-infected 3T3 cells is likely due to a strong reduction in plaque size, with only the largest plaques remaining visible. We have rephrased the relevant section to clarify this point and avoid any potential confusion:

      “In agreement with our previous results, only small plaques were formed in 3T3 cells upon infection with the ΔgcvPAB mutant (plaque area: 15±19% of wild-type level) and small plaques were also formed by the complemented strain in the absence of IPTG (50±19%).”

      (2) Do other Listeria strains contain the stop codon in fhs? How common is this mutation? That would be interesting to know.

      We determined the frequency of inactivated fhs genes among 30,000 publicly available L. monocytogenes genomes. The analysis identified only 10 isolates carrying truncated fhs alleles. These isolates fell into two groups: (i) EGD-e and its descendants, and (ii) a cluster of five ST2 food isolates. These findings have been added as a separate results section.

      (3) Based on the observation that fhs+ ΔgcvPAB ΔglyA mutant is only possible to isolate in complex media, and fhs is responsible for converting formate to 1C-THF with the addition of FolD, have the authors thought of supplementing synthetic media with formate and assessing mutant growth?

      No, we did not test formate supplementation. However, we included additional experiments testing adenine and thymine supplementation (Fig. 6E and 7B). These results show that purine and thymine become limiting in mutants lacking 1C-THF synthesizing pathways.

      Reviewer #3 (Public review):

      Summary:

      In this study, Freier et al. demonstrate that 3 distinct metabolic pathways are critical for the synthesis of 1C-THF, a metabolite that is crucial for the growth and virulence of Listeria monocytogenes. Using an elegant suppressor screen, they also demonstrate the hierarchical importance of these metabolic pathways with respect to the biosynthesis of 1C-THF.

      Strengths:

      This study uses elegant bacterial genetics to confirm that 3 distinct metabolic pathways are critical for 1CTHF synthesis in L. monocytogenes, and the lack of either one of these pathways compromises bacterial growth and virulence. The study uses a combination of in vitro growth assays, macrophage-CFU assays, and murine infection models to demonstrate this.

      Weaknesses:

      (1) The primary finding of the study is that the perturbation of any of the 3 metabolic pathways important for the synthesis of 1C-THF results in reduced growth and virulence of L. monocytogenes. However, there is no evidence demonstrating the levels of 1C-THF in the various knockouts and suppressor mutants used in this study. It is important to measure the levels of this metabolite (ideally using mass spectrometry) in the various knockouts and suppressor mutants, to provide strong causality.

      As already outlined above, we do not have any experimental possibilities to measure 1C-substituted THF in L. monocytogenes extracts directly. However, to provide additional evidence for our interpretation that “Three pathways feed the 1C-THF pool…”, we included additional genetic experiments.

      The first experiment demonstrates that the growth defect of the fhs- ΔglyA igcvPAB strain in synthetic medium lacking IPTG—which reflects the synthetic lethality of fhs with glyA and gcvPAB—can be rescued by the addition of adenine (Fig. 6E). This indicates that the fhs/fold pathway, GlyA, and the glycine cleavage system are essential due to their combined contribution to purine biosynthesis. This result confirms the metabolic model presented in Fig. 1A and thus supports the hypothesis that all three pathways contribute to 1C-THF biosynthesis.

      The second experiment additionally demonstrates synthetic lethality of gcvPAB and glyA with the fold gene. FolD acts downstream of Fhs and is one of the three enzymes synthesizing N5, N10methylene-THF shown in Fig. 1A. The fold gene is essential in EGD-e (PMID: 36114002), most likely explained by fhs inactivation. However, we were able to delete fold in an EGD-e background carrying a reconstituted fhs gene and the resulting fhs<sup>+</sup> Δfold strain was as viable as a fhs<sup>+</sup> ΔglyA ΔgcvPAB strain (Fig. 7A). However, a fhs<sup>+</sup> Δfold ΔglyA igcvPAB strain required IPTG for growth in BHI medium (Fig. 7A), indicating that the simultaneous deletion of glyA and gcvPAB is not tolerated in the absence of fold, similar to what is observed in the absence of fhs. Notably, this growth defect was not rescued by adenine supplementation (Fig. 7B), but was alleviated by thymine addition, which is also consistent with the metabolic model shown in Fig. 1A.

      Even though we are unable to directly demonstrate reduced 1C-THF levels, we hope that these two genetic approaches together with the revised title and heading of the relevant paragraph will, in the reviewers’ eyes, support our hypotheses.

      (2) The story becomes a little hard to follow since macrophage-CFU assays and murine infection model data precede the in vitro growth assays. The manuscript would benefit from a reorganization of Figures 2,3, and 4 for better readability and flow of data.

      We respectfully disagree with the reviewer. The attenuation of the ΔgcvPAB mutant in macrophages and fibroblasts was the primary motivation for further investigating its phenotype. Therefore, we chose to begin the manuscript with results from various virulence studies before presenting the sections that provide mechanistic explanations. In our view, this sequence represents a more logical and coherent way to present the data.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Synthetic medium assumptions: LSM "mimics" intracellular limitation but isn't chemically validated against host cytosolic composition. Nutrient availability conclusions could be biased.

      This is correct. We added this information:

      “LSM broth is a chemically defined medium that contains all components required for growth at defined concentrations, but it has not been chemically validated to reflect host cytosolic conditions…”

      (2) Glycine toxicity: The paper introduces the concept of glycine toxicity in ΔgcvPAB mutants. However, the conditions under which glycine becomes toxic could be further elucidated. Why glycine causes toxicity despite being an essential metabolite in other contexts requires a more in-depth mechanistic explanation.

      The concept of glycine toxicity in GCS mutants has been described previously by other researchers. We have added further details to better explain glycine toxicity and how it accounts for the growth phenotype of the ΔgcvPAB mutant:

      “If glycine cannot be catabolized (and 1C-THF cannot be generated) by the GCS due to deletion of gcvPAB, glycine might be re-routed to the serine hydroxymethyl transferase GlyA for serine formation, even though this would consume 1C-THF and therefore even further deplete the cell for 1C-THF.” and

      “In the complete absence of glycine, growth of the ΔgcvPAB mutant was largely unaffected, presumably because glycine cannot be converted to serine by GlyA anymore, thereby conserving the 1C-THF pool.”

      (3) Figures:

      (a) Figure 1B, scale bar?

      A scale bar was added.

      (b) Figure 1C, the standard error bars for igcvPAB (both with and without IPTG) are relatively wide, indicating high variability in the data. This suggests that the results in the igcvPAB group are not as consistent as the wild-type (wt) or ΔgcvPAB groups. Please show the original data and perform a statistical test (e.g., t-test or ANOVA).

      The original data have been added to Fig. 1B. t-test results (with Bonferroni-Holm correction) are now included for all samples.

      (c) Figure 2D, scale bar?

      These are sections of agar plates. From our point of view, a scale bar does not add relevant information.

      (d) Figure 3, please check the labels of the Figure 3 legend. (D) and (E) or A-B?

      Thanks, corrected.

      (e) Figure 5E, quantification of plaque areas?

      The plaque areas were quantified. A blot showing these quantitative data is now presented in Fig. 5F.

      Reviewer #2 (Recommendations for the authors):

      Line 58: There are published studies that indicate that syncytiotrophoblasts are actually resistant to Listeria infection and that it is extravillous trophoblasts that are likely to serve as entry points for Listeria into the placenta (see, for example, Lowe et al, Infect Immun. 2018 Volume 86 Issue 6 e00801-17).

      We thank the reviewer for this comment and have removed our statement claiming that syncytiothrophoblasts are the entry point as this is not relevant to the understanding of the work presented here.

      Reviewer #3 (Recommendations for the authors):

      (1) Please mention the number of times experiments were performed as independent biological replicates, wherever applicable.

      We added this information to the figure legends wherever this was necessary.

      (2) Please provide the details of the type of statistical analysis used for the various graphs, either in the figure legends or in the materials & methods section.

      This information was also added to the figure legends wherever it still was missing.

      (3) Can the authors comment on how the weight-loss phenotype of animals and the variation in the size of the spleen between animals infected with wild-type and mutants in Figure 2 can be explained without any significant changes in the CFU? Additionally, I did not see details regarding the number of animals used in the murine infection model experiments. Please mention this along with the type of statistical analyses used.

      We do not see a contradiction here, as the apparent differences are explained by the distinct time points at which CFU numbers (day 3 post-infection) and spleen size (day 9 postinfection) were measured. Starting from day 8, the difference in weight between animals infected with the wild-type strain and those infected with the ΔgcvPAB mutant becomes clear for the first time. At day 3, no significant differences are detected in either CFU numbers or weight. By day 9, when the weight difference has become apparent, differences in spleen size are also observed. To improve clarity, the time points at which each analysis was performed have been added to Fig. 2G and Fig. 2I. The number of infected animals and the type of statistical analysis used are now specified in the figure legend.

    1. eLife Assessment

      This is a detailed and well-designed simulation study of the utility of replication metrics in animal-to-human study translations in bridging the gap between laboratory discoveries and health practice, a critical consideration in turning laboratory scientific research findings into tangible, real-world applications, to directly help human health. The study approaches are convincing, and the findings are important, as they offer insights into clinical research translations to advance health decision-making.

    2. Reviewer #1 (Public review):

      [Editors' note: This revised version of your article has been assessed by the Reviewing Editor without further input from the original reviewers. The comments raised by the original reviewers in the earlier round of review have been addressed. The study findings are quite insightful and important, and the evidence is strong, convincing, and a substantial addition to the evidence base.]

      A well-designed and preregistered simulation study investigating whether replication-success metrics can be applied to assess animal-to-human translation. The study is comprehensive, uses realistic parameter settings, and provides valuable insights into how different metrics behave under varied conditions.

      Strengths:

      (1) Methodologically rigorous and transparently preregistered.

      (2) Comprehensive simulation design covering a wide range of plausible scenarios.

      (3) Clear description of metrics and decision rules.

      (4) Valuable contribution to understanding the limitations of applying replication metrics to translation questions.

    3. Reviewer #2 (Public review):

      Summary:

      The authors attempt to address the issue of high rates of translation failure from animal studies to humans in the literature, where promising results in animal studies fail when conducting human clinical trials. Using parameters from a previous meta-analysis on prenatal amino acid supplementation and the effects it has on maternal blood pressure, the authors assessed the performance of the metrics used and whether they can quantify translation success. Performing a simulation study, the authors compared nine translation success metrics and found that no one method was uniformly optimal. The authors list several limitations of the study, such as comparability of effect sizes between animal and human studies, different goals of animal studies versus human studies, and the focus of the study on one aspect (statistics of translation) is part of a broader, more complex decision-making process before proceeding to human trials. The authors recommend using multiple metrics in combination while taking into consideration their strengths and weaknesses to assess the translation of animal studies to human outcomes. The paper achieves the aim of providing a model with several metrics to evaluate translation success from animal studies to humans.

      Strengths:

      (1) Utilizing 9 different translation success metrics in combination provides strong flexibility in evaluating whether results in animal studies can translate to humans. This would allow researchers to evaluate translation success using multiple different metrics according to the context of the study.

      (2) The authors accommodated for the limited sample size in animal studies, which are typically underpowered, and also caution that special attention should be given to heterogeneity when interpreting translation results.

      (3) Overall, this approach has the potential to be applied to other biomedical studies, provided the limitations for each of the metrics are considered. It would provide a useful tool in assessing translation from animals to humans, in addition to other factors such as safety, pharmacokinetics, etc.

      Weaknesses:

      While the study has several strengths, there are some limitations.

      (1) Preclinical animal study sizes tend to be much smaller than human studies, which results in underpowered results. The authors adjusted for this by pooling animal study data. However, high heterogeneity in the animal studies can affect translation results.

      (2) The study focuses only on evaluating the statistical component of translation, which is only one aspect of the decision-making process to move on to human trials. The study does not take into account safety and toxicological profiles, pharmacokinetics, or genetics, which are important considerations that influence the overall effect in humans.

    4. Reviewer #3 (Public review):

      Summary:

      This paper focused on how to navigate the complex decision-making process of whether to go into human trials. This is a critical topic considering the well-documented challenges in replicating and translating findings. While these are two distinct topics (i.e., replication and translation), they are related, and the authors simulated many conditions to assess the utility of replication assessment metrics.

      Strengths:

      A major strength of the study is the detailed approach to identifying relevant conditions and metrics, and to providing rich results that outline the strengths and weaknesses of each metric. Any simulation study is challenged by trying to identify the most relevant variables of interest, and this study provided sound justification for its chosen variables of interest. While this study does not make a strong recommendation (which I see as a strength), it does provide a comprehensive overview of the various metrics and conditions that were investigated.

      Conclusion:

      This paper provides a much-needed investigation and discussion of how decisions are made when assessing whether to go into human trials. This is an important topic that productively challenges the status quo, considering documented challenges in replication and translation in biomedical research.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      A well-designed and preregistered simulation study investigating whether replication-success metrics can be applied to assess animal-to-human translation. The study is comprehensive, uses realistic parameter settings, and provides valuable insights into how different metrics behave under varied conditions.

      Strengths:

      (1) Methodologically rigorous and transparently preregistered.

      (2) Comprehensive simulation design covering a wide range of plausible scenarios.

      (3) Clear description of metrics and decision rules.

      (4) Valuable contribution to understanding the limitations of applying replication metrics to translation questions.

      Weaknesses:

      (1) The conceptual distinction between replication and translation could be more clearly emphasized.

      (2) Interpretation of results is dense and can be challenging to follow without a clear and summarized.

      (3) Some simulation parameters (effect sizes, heterogeneity, and number of animal studies) require more substantial justification.

      (4) Practical recommendations could be more explicit to guide applied researchers.

      We thank Reviewer 1 for the general positive assessment of our study and for the constructive feedback. We have addressed all of the four identified weaknesses in the revised manuscript. Specifically,

      (1) Conceptual distinction between replication and translation. We have reinforced this distinction at multiple points in the manuscript: in Section 2.7 (just before introducing the translation success metrics), in the Discussion, and in a new working definition of translation success added to the Introduction. We further explicitly acknowledge that statistical translation success, as defined here, is narrower than biological translation.

      (2) The dense result section. We have added a summary Table (Table 3) at the end of the Results section that compares all metrics on key properties (strengths and weaknesses, overall type 1 error control, sensitivity to heterogeneity, dependence on animal sample size and number of studies, and behaviour under effect mismatch). We also direct readers to this table early in Section 3.2, so that readers less interested in the technical details can obtain the key take-home messages without reading the full section.

      (3) Further justification of simulation parameters. We have substantially extended the rationale for our parameter choices in Section 2.4 and the Limitations section. We explain that our parameters are grounded in an empirical meta-analytic dataset, contextualise the large effect size and heterogeneity value against published benchmarks from preclinical research, and clarify that our main goal was to explore directional trends rather than absolute performance under specific values. We have also added an invitation for others to explore alternative parameter spaces using our openly available code.

      (4) Practical recommendations. We have extended the Recommendations section (pages 21–22) with more explicit scenario-specific guidance, supported by the new summary table.

      Reviewer #2 (Public review):

      Summary:

      The authors attempt to address the issue of high rates of translation failure from animal studies to humans in the literature, where promising results in animal studies fail when conducting human clinical trials. Using parameters from a previous meta-analysis on prenatal amino acid supplementation and the effects it has on maternal blood pressure, the authors assessed the performance of the metrics used and whether they can quantify translation success. Performing a simulation study, the authors compared nine translation success metrics and found that no one method was uniformly optimal. The authors list several limitations of the study, such as comparability of effect sizes between animal and human studies, different goals of animal studies versus human studies, and the focus of the study on one aspect (statistics of translation) is part of a broader, more complex decision-making process before proceeding to human trials. The authors recommend using multiple metrics in combination while taking into consideration their strengths and weaknesses to assess the translation of animal studies to human outcomes. The paper achieves the aim of providing a model with several metrics to evaluate translation success from animal studies to humans.

      Strengths:

      (1) Utilizing 9 different translation success metrics in combination provides strong flexibility in evaluating whether results in animal studies can translate to humans. This would allow researchers to evaluate translation success using multiple different metrics according to the context of the study.

      (2) The authors accommodate for the limited sample size in animal studies, which are typically underpowered, and also caution that special attention should be given to heterogeneity when interpreting translation results.

      (3) Overall, this approach has the potential to be applied to other biomedical studies, provided the limitations for each of the metrics are considered. It would provide a useful tool in assessing translation from animals to humans, in addition to other factors such as safety, pharmacokinetics, etc.

      Weaknesses:

      While the study has several strengths, there are some limitations.

      (1) Preclinical animal study sizes tend to be much smaller than human studies, which results in underpowered results. The authors adjusted for this by pooling animal study data. However, high heterogeneity in the animal studies can affect translation results.

      (2) The study focuses only on evaluating the statistical component of translation, which is only one aspect of the decision-making process to move on to human trials. The study does not take into account safety and toxicological profiles, pharmacokinetics, or genetics, which are important considerations that influence the overall effect in humans.

      We thank Reviewer 2 for the thoughtful summary and for recognising the strengths of our study. We believe that both weaknesses were addressed in the revised version of our manuscript. Specifically,

      (1) Heterogeneity in animal studies. We agree that high heterogeneity in animal studies is an important limitation, and we address it directly in our simulation design by including a wide range of heterogeneity values (including very high levels, as observed in animal studies). Our results show clearly how heterogeneity affects the performance of each metric, and we highlight this in both the new summary Table (Table 3) and the Recommendations section which was extended. We also caution applied researchers to pay special attention to heterogeneity when interpreting translation results.

      (2) Focus on the statistical component of translation. We fully agree that statistical translation success is only one aspect of a broader decision-making process. We have elaborated on this in the revised manuscript, both in a new working definition of translation success in the Introduction (which explicitly distinguishes statistical from biological translation) and in a new paragraph in the Discussion section where we situate our metrics within translational decision-making frameworks such as PATH. They make it clear that progression to human trials depends on a suite of evidence of which statistical translation is only one part.

      Reviewer #3 (Public review):

      Summary:

      This paper focused on how to navigate the complex decision-making process of whether to go into human trials. This is a critical topic considering the well-documented challenges in replicating and translating findings. While these are two distinct topics (i.e., replication and translation), they are related, and the authors simulated many conditions to assess the utility of replication assessment metrics.

      Strengths:

      A major strength of the study is the detailed approach to identifying relevant conditions and metrics, and to providing rich results that outline the strengths and weaknesses of each metric. Any simulation study is challenged by trying to identify the most relevant variables of interest, and this study provided sound justification for its chosen variables of interest. While this study does not make a strong recommendation (which I see as a strength), it does provide a comprehensive overview of the various metrics and conditions that were investigated.

      Weaknesses:

      The weaknesses of the study are the limited focus on specific metrics, the assumptions, particularly in the limited number of human study variables, and the less-than-ideal approachable summary of findings for a non-technical audience.

      Conclusion:

      This paper provides a much-needed investigation and discussion of how decisions are made when assessing whether to go into human trials. This is an important topic that productively challenges the status quo, considering documented challenges in replication and translation in biomedical research.

      We thank Reviewer 3 for the positive assessment and for the constructive suggestions.

      We have addressed the identified weaknesses as follows:

      (1) The assumptions around human study variables. We acknowledge these as inherent constraints of the simulation design. We have added a note in the Limitations section about the fixed human sample size (N = 107 per group), clarifying that while this value is grounded in a power analysis as per regulatory standards, it represents one particular scenario and may not generalise to all contexts. Further, we have contextualised and motivated the other simulation parameters better. We also invite readers to explore alternative conditions using our openly available code.

      (2) Approachability of the summary of findings for a non-technical audience. We have added a summary Table (Table 3) at the end of the Results section, comparing the metrics on key properties including overall type 1 error control, sensitivity to heterogeneity, dependence on animal sample size and number of studies, and behaviour under effect mismatch. We direct readers to this table early in Section 3.2 so that those less interested in the technical details can obtain the main take-home messages without reading the full section.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major points:

      (1) Conceptual framing: clearer distinction between replication vs translation

      The Introduction correctly points out the conceptual difference between replication and translation (animal to human), but this distinction needs to be reinforced repeatedly, especially when interpreting metric performance. For instance, several metrics (e.g., meta-analysis, replication BF) inherently assume exchangeability of findings, which is rarely justified in translation because species differ biologically.

      The manuscript should explicitly state why treating animal findings as "original studies" and human findings as "replications" can be misleading. Add a subsection in the Discussion: Why replication metrics behave differently in translation settings. This will help provide a more straightforward interpretation of the results beyond the numerical findings.

      Thank you for your feedback. While the purpose of our study is to assess the applicability of the replication success metrics in the translation context, we agree that the reader should be reminded that these two concepts differ and metrics’ assumptions might not always hold. We have reiterated the difference between replication and translation in Section 2.7, just before we introduce the translation success metrics (see end of page 7). We also reiterate it in the Discussion (see end of page 20). Here, we emphasise that while some of the metrics assume that both studies investigate the same effect, this is unlikely to be the case in translation, leading to some of the metric’s assumptions being violated which impacts the performance of the metrics.

      (2) Stronger justification of the simulation parameters is needed

      The simulation factors are comprehensively presented (Table 1), but certain choices appear arbitrary or oversimplified.

      - Effect sizes: The three levels (0, −4.44, −24.37) are derived from the motivating dataset, but the paper should explain that these represent extremely large effects in many biomedical contexts.

      - Heterogeneity values: τ<sup>2</sup> = 291.1 is enormous; adding context about real-world heterogeneity distributions would help.

      - Number of animal studies (k): Using only 2-5 studies may not reflect reality; many preclinical fields have >30 studies before clinical translation.

      These choices should be more explicitly defended in Section 4 (Limitations), beyond the brief mention already there. Provide a sensitivity analysis, or explain why extrapolation beyond this parameter space is reasonable.

      We agree that the choice of the parameter values might sometimes appear arbitrary. However, instead of arbitrarily choosing parameter values, we base our choice on data from a meta-analysis. This particular meta-analysis might not be representative of all of pre-clinical and clinical research, but because Terstappen included both animal and human studies investigating the same research question it was particularly well suited. They further used an outcome (maternal blood pressure) that is comparable between rats and humans, which is quite rare. We have specified this further in Section 2.4 (Motivating dataset, page 5). In the Limitations section, we acknowledge any possibly unrealistic simulation conditions again, and emphasize that our main goal was to explore trends in the metrics’ behavior as the conditions changed rather than their absolute performance under specific values. Further, the effect sizes (0, −4.44, −24.37 mmHg) span a meaningful range on the unstandardized mean difference scale for blood pressure measurements: from no effect to a modest but clinically relevant reduction to a large effect typical of animal studies. The large heterogeneity value corresponds to a relative heterogeneity of I^2 of 95.33% in the animal meta-analysis. While this appears high, it is frequently observed in preclinical research: Hooijmans et al (2022) showed that 55% of animal study meta-analyses using mean differences as effect size measure have I^2>75%. We also added a footnote reiterating the fact that such high effect sizes (on the raw mean difference scale) are indeed common in animal studies (on page 6). Regarding k, we acknowledge that pooling only 2 to 5 animal studies may not reflect common practice. However, the directional trends in type 1 error and power are clearly visible in our Figures. Larger k decreases the type 1 error of the animal studies, while the power is increased unless there is high heterogeneity between animal studies and there is only a small effect. Extending the range further is unlikely to change the conclusions. Moreover, in practice, the decision to advance to human trials considers evidence well beyond the statistical considerations we simulate. All of the above is now emphasized more explicitly in both the methods, where we have substantially extended the reasoning for choosing the simulation conditions, and the limitations section. Finally, we added an invitation to others to use our open material (i.e., code) and explore the behaviour of the metrics under other conditions (see top of page 21).

      (3) Decision criteria (strict/lenient/no criterion) need a clearer rationale

      The three continuation rules are a strength of the study, but:

      - The lenient criterion (any negative estimate is considered "beneficial") is unrealistic and should be reframed.

      - The strict criterion (p < 0.025) heavily inflates effect sizes (in Figure 1b) and may distort interpretation.

      It would be helpful to provide a table showing, for each criterion, its real-world analogue (e.g., regulatory requirement, exploratory progression, mechanistic plausibility).

      We have followed your suggestion and added a Table (Table 2) with the description of the criterion and a description of its real-world analogue. No criterion represents an important reference scenario used to evaluate metric behaviour independent of progression decisions. The strict criterion is the closest to regulatory-style evidence. It is also highly selective and therefore might induce biases (e.g., inflated effect sizes). We link lenient to an exploratory decision-making where efficacy evidence is considered in addition to other factors (e.g., safety), but not intended to represent a certain regulatory standard.

      (4) Interpretation of simulation results needs more focus

      The Results section is extremely detailed, making it challenging to identify the central take-home messages. The authors should consider adding a concise summary table comparing metrics on key properties:

      - T1E control robustness.

      - Sensitivity to heterogeneity.

      - Dependence on animal sample size.

      - Dependence on k.

      - Bias under asymmetric effects.

      Moving some nested-loop plot descriptions to the Supplement. Right now, descriptions are technically correct but cognitively heavy.

      We agree with your comment and have attempted to implement it in our summary Table 3, at the end of the results section. We also point readers early on to the Table, so that they can skip the more technical and detailed description if they want (see first paragraph section 3.2, page 12). After some trial and error, we agreed that the chosen columns are the most useful for an applied researcher to get a quick overview. Our table now summarises for each metric its main strengths and weaknesses, its behaviour with increasing heterogeneity, its sensitivity to more animal data (i.e., larger k and larger animal sample size), and its behaviour under effect mismatch (i.e., when the true effect in the animal and human study are dissimilar).

      (5) The discussion should provide explicit recommendations.

      The authors provide high-level recommendations, but the recommendations lack specific guidance. When heterogeneity is low, controlled sceptical p-value works well. When effect sizes differ: weighted Edgington is stable. The authors should avoid using replication BF when the animal effect ≠ human effect. Meta-analysis should not be used when human heterogeneity is high, because of inflated T1E.

      We agree that explicit recommendations would be helpful to the applied researcher. As mentioned in the reply to the previous comment, we have added a summary table which lists the strengths and weaknesses of each metric. We also extended the paragraph in the Recommendations section (on page 21 and 22) to give some examples of scenarios in which certain metrics would be recommended over others.

      (6) Recommendations for applied researchers

      The study is missing an explicit definition of "translation success". The manuscript implicitly defines translation success as: "Both animal and human results show a beneficial treatment effect according to metric X". But this is different from biological translation, which concerns underlying mechanisms. The authors briefly mention this conceptual challenge, but this should be elaborated, as it is central to interpretation.

      Thank you for this comment. We agree that “translation success” was not explicitly defined. We have now added a working definition in the Introduction, clarifying that, in this paper, translation success is defined statistically, and depends on the metric. We now explicitly acknowledge that this is a narrower definition than biological translation. We also elaborate on this distinction in the Discussion where we note that the appropriate metric and interpretation of translation success depends on the translation goal and that statistical translation is distinct from biological translation.

      Minor points:

      (1) The abstract could include a direct sentence on the main conclusion. For example, no metric was uniformly optimal; controlled sceptical p-value and weighted Edgington performed most consistently.

      Our abstract already included main conclusions. We added the word “However” to emphasize the sentence “no metric was uniformly optimal” a bit more.

      (2) The figures are informative, but nested loop plots are very dense. Consider providing a guided example in the figure caption explaining how to read them (as partially done in Figure 1a, but repeat for all).

      We agree that the Figures can be very overwhelming at first. We did not want to add specific helping elements as we did in Figure 1 to not make the figures even busier. The goal was to introduce the reader gently to the nested loop plots via Figure 1 before having them look at the remaining figures. We hope that with the added summary Table and the more detailed recommendations, applied researchers less interested in the statistical details will still find the information most relevant for them easily.

      (3) Methods: Section 2.4 could clearly state that effect sizes are in units of mmHg (blood pressure) from the dataset.

      Thank you for pointing this out. This has been added.

      (4) Results: This section is long; consider adding a brief summary paragraph at the end of 3.2.

      We added a summary table, allowing interested readers to skip the long section entirely.

      (5) Limitations: Add a note about publication bias in animal studies (you mention it in the Introduction, but not in Limitations). Add a statement about effect direction consistency (i.e., animal effect negative but human positive), which is not explored in the simulation grid.

      Thank you for pointing out this inconsistency. A note about publication bias in animal studies was added to the Limitations section (that this was not investigated). A note about opposite animal and human effects was added to Section 2.5 (Simulation conditions) under “Animal and human effect sizes”.

      Reviewer #2 (Recommendations for the authors):

      Animal studies are typically highly controlled, using animal models that are either outbred to provide higher genetic variability or inbred with very little genetic variability and with a specific phenotype. Additionally, many rodent models are incomplete models of the overall human phenotype and are typically used to investigate only one aspect of the condition/disease. Some of the rat animal models that the Terstappen et al. (2020) systematic review used as the simulation parameters for the study included outbred (Sprague-Dawley, Wistar) and inbred Spontaneous Hypertensive Rats (SHR), which have different mechanisms in which hypertensive onset can occur, especially if inducing preeclampsia in outbred animals. Is it feasible to reduce heterogeneity in the animal results if only outbred or only SHR are considered instead? I realize this may reduce the sample size even further.

      You raise an important point differentiating biological (rather than statistical) translation. We have added a sentence about differences between rat models and humans to the new paragraph in the Limitations section (bottom page 20 and top page 21) on the distinction between biological and statistical translation. As for reducing heterogeneity in the animal results by focusing on one type of rats, we agree focusing on one type of rats might reduce heterogeneity. We however consider this reduction to be very small (because the results of the study with SHR are actually comparable to the results with Wistar and SD rats). Therefore, rerunning the simulation would not yield results that differ in any meaningful way from those already reported and the substantial computational effort required to do so is not warranted.

      Reviewer #3 (Recommendations for the authors):

      Overall, I found this a very detailed study. However, my recommendation is to provide a more approachable overview of the results to reach a wider audience. Currently, the article is much more technical and statistically focused. I think two additions could help.

      (1) A summary table of each of the metrics and their strengths and weaknesses under the various conditions (e.g., animal and human study characteristics). Currently, this is done via text, but I think a high-level summary via a table could be a compelling way to make the simulations more approachable for a non-technical audience.

      As requested also by reviewer 1, we have added a summary table.

      (2) Contextualize the findings within the decision-making process a little more. The authors have a well-written limitations section that acknowledges this; however, I think the discussion (and maybe the introduction) could be enriched by putting the simulation findings into context. For example, this paper suggests a framework that includes replication as part of the decision-making process for human trials (https://www.cell.com/med/fulltext/S2666-6340(24)00296-4).

      We agree that situating our metrics within existing translational decision-making frameworks adds important context. We have added a paragraph in the Discussion (before the Limitations section on page 21) clarifying that the metrics evaluated here should not be viewed as standalone decision rules for progression from animal studies to human trials. Several frameworks have recently emerged precisely to guide such decisions in a more structured, multidimensional way. We refer to PATH and also to the GALENOS approach [DOI: 10.1186/s12874-026-02891-4]. Within such frameworks, translation success metrics of the kind evaluated here may provide a quantitative assessment of the consistency between animal and human efficacy findings, thereby informing one component of a broader translational evidence assessment. We have also briefly mentioned at the end of the Introduction (page 4) that frameworks for structuring the use of preclinical evidence in translational decisions are being developed, further motivating the need for quantitative tools such as those evaluated here.

      Below are some additional minor comments for the authors to consider:

      (1) In the abstract (4th line), there is an extra 'l' in failure.

      Thank you for the detailed review. We have fixed this.

      (2) I think since the study is completed, the objectives in the introduction should be past tense, not future.

      We have fixed this.

      (3) The limitations section should include the fixed human sample size. N=107 per group is grounded in the literature, but this varies widely based on the effect size of interest. Again, not material to the point of translation under simulated conditions (of which this would have increased the simulations well above the 648 already included), but given the impact this has on insights, this limits this investigation to a degree and should be acknowledged.

      Thank you for your comment. We have added a note about the human sample size to the paragraph about the simulation conditions in the Limitations section. The human sample size was computed via power analysis as per regulations, but we realize this could change depending on the effect size.

      (4) I appreciate how shrinkage was calculated. Though it is worth noting that the Reproducibility Project: Cancer Biology found much higher rates, which are similar to reports from biotech and pharma (e.g., 11% and 20-25% for Amgen and Bayer).

      We already mentioned the high rates of shrinkage in the Replication Project Cancer Biology (see page 10). We have now also emphasised that one could adapt these levels further depending on the situation.

    1. eLife Assessment

      This fundamental study uses simultaneous EEG and fMRI recordings to shed light on the relationship between alpha and gamma oscillations and specific cortical layers. The sophisticated methodology provides compelling evidence for correlations between oscillatory power and the strength and contents of fMRI signals in different cortical layers. This paper will be of interest to neuroscientists studying the role and mechanisms of alpha and gamma oscillations.

    2. Reviewer #1 (Public review):

      In this manuscript, Clausner and colleagues use simultaneous EEG and fMRI recordings to clarify how visual brain rhythms emerge across layers of early visual cortex. They report that gamma activity correlates positively with feature-specific fMRI signals in superficial and deep layers. By contrast, alpha activity generally correlated negatively with fMRI signals, with two a higher frequency within the alpha reflecting feature-specific fMRI signals. This feature-specific alpha code indicates an active role of alpha oscillations in visual feature coding, providing compelling evidence that the functions of alpha oscillations go beyond cortical idling or feature-unspecific suppression.

      The study is very interesting and timely. Methodologically, it is state of the art. The findings on a more active role of alpha activity that goes beyond the classical idling or suppression accounts is in line with recent findings and theories. In sum, this paper makes a very nice contribution to the literature. In particular, it provides a novel characterization of how oscillatory signals orchestrate the coding of visual contents in the visual cortex and provides a starting point for further research examining how this oscillatory coding changes across visual contents and tasks.

    3. Reviewer #2 (Public review):

      The authors address a long-standing controversy regarding the functional role of neural oscillations in cortical computations and layer-specific signalling. Several studies have implicated gamma oscillations in bottom-up processing, while lower-frequency oscillations have been associated with top-down signalling. Therefore, the question the authors investigate is both timely and theoretically relevant, contributing to our understanding of feedforward and feedback communication in the brain. This paper presents a novel and complicated data acquisition technique, the application of simultaneous EEG and fMRI, to benefit from both temporal and spatial resolution. A sophisticated data analysis method was executed in order to understand the underlying neural activity during a visual oddball task. The authors defined both feature-specific and feature-unspecific contrasts, further subdivided by EEG power regressors, to examine how orientation information is signalled across cortical layers. Feature specific contrast was established via comparing trials where stimulus orientation (respectively) was left with those where the stimulus orientation was right. Further specifying it depending on EEG power regressors as congruent where stimulus orientation of EEG regressor matches voxel preference or incongruent (stimulus orientation of EEG regressor does not match voxel preference).

      Figures are well-designed and appropriately represent the results, which seem to support the overall conclusions. However, some of the claims (particularly those regarding the contribution of gamma oscillations) feel somewhat overstated, as the results offer indeed some significant evidence. On the other hand, the lower-frequency findings are compelling, the functional specificity observed within the alpha frequency band is a particularly interesting result and further highlights the importance of distinguishing feature specificity in order to reveal more nuanced characteristics of neuroimaging data.

      Overall, main findings are very interesting, and mainly in line with our current understanding of feedback and feedforward signalling. The paper is well-written, addresses a relevant and timely research question, introduces a novel and elegant analysis approach, and presents interesting findings.

      The evidence for gamma involvement in the observed effects is selective: no significant gamma-related clusters were found for the feature-unspecific BOLD signal (Figure 5C,F), with significant effects emerging only in positively responding voxels and only for the contrast between congruent and incongruent conditions in the feature-specific BOLD response. The authors address this in the Discussion, noting that the stimulus may have elicited a weaker gamma response overall, and the contrast of EEG congruent vs. incongruent is necessary in order to achieve the largest contrast-to-noise ratio.

      Authors reported negative relationship between the alpha frequency band and the feature specific BOLD signal increases (for congruent condition, Figure 5A,D). Furthermore, testing for the functional specificity between lower vs. upper alpha (Figure 5B,E), the authors included statistical test on the mixed effects model coefficients, and found significant interaction between alpha frequencies in the congruent condition. This interaction was mainly driven by the upper alpha band (which was later confirmed with simple effects analysis) and revealed stronger negative relationship of upper alpha and the BOLD signal for the subtraction of congruent over incongruent conditions. These are exciting results, which further advocate for a more active role of upper alpha band involvement (relative to lower alpha band) in processing visual features.

      Expanding on this, the authors have also conducted an exploratory analysis of the relationship between the behavioural findings and underlying neural activity for non-oddball trials (Figure S12 in Supplementary Figures). This confirmed a positive relationship between task performance and alpha frequency, suggesting that high behavioural accuracy is reflected by a stronger modulation of high-frequency alpha power.

      This study provides a valuable and exciting contribution to the literature on oscillatory dynamics and laminar fMRI.

      Comments on revised version.

      Thank you for the thorough revision and for addressing the comments so carefully. The new figures are super beautiful and make the results considerably easier to interpret, they are a real improvement to the paper.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In this manuscript, Clausner and colleagues use simultaneous EEG and fMRI recordings to clarify how visual brain rhythms emerge across layers of early visual cortex. They report that gamma activity correlates positively with feature-specific fMRI signals in superficial and deep layers. By contrast, alpha activity generally correlated negatively with fMRI signals, with two higher frequencies within the alpha reflecting feature-specific fMRI signals. This feature-specific alpha code indicates an active role of alpha oscillations in visual feature coding, providing compelling evidence that the functions of alpha oscillations go beyond cortical idling or feature-unspecific suppression.

      The study is very interesting and timely. Methodologically, it is state-of-the-art. The findings on a more active role of alpha activity that goes beyond the classical idling or suppression accounts are in line with recent findings and theories. In sum, this paper makes a very nice contribution. I still have a few comments that I outline below, regarding the data visualization, some methodological aspects, and a couple of theoretical points.

      The authors put a lot of effort into the figure design. For instance, I really like Figure 1, which conveys a lot of information in a nice way. Figures 3 and 4, however, seem over engineered, and it takes a lot of time to distill the contents from them. The fact that they have a supplementary figure explaining the composition of these figures already indicates that the authors realized this is not particularly intuitive. First of all, the ordering of the conditions is not really intuitive. Second, the indication of significance through saturation does not really work; I have a hard time discerning the more and less saturated colors. And finally, the white dots do not really help either. I don't fully understand why they are placed where they are placed (e.g., in Figure 3). My suggestion would be to get rid of one of the factors (I think the voxel selection threshold could go: the authors could run with one of the stricter ones, and the rest could go into the supplement?) and then turn this into a few line plots. That would be so much easier to digest.

      We thank the reviewer for their insightful comments. Below we will address each point separately and highlight the changes made to the manuscript. In agreement with the reviewer we have recompiled Figures 4 and 5 (previously Figures 3 and 4). The new figures only present results for the 10% voxel selection threshold (with 5% and 25% moved to Supplementary Figures, see Figures S1-S9). Instead of the radially arranged layout, we opted for a more traditional figure layout, which significantly improved readability.

      (2) The division between high- and low-frequency alpha in the feature-specific signal correspondence is very interesting. I am wondering whether there is an opposite effect in the feature-unspecific signal correspondence. Would the high-frequency alpha show less of a feature-unspecific correlation with the BOLD?

      Following the reviewer’s interesting suggestion, we added the low/high frequency alpha analysis to the feature-unspecific analysis. Indeed, we have found a significant interaction between the sign of the signal change for selected voxel (positive vs. negative BOLD) and alpha sub-band (low vs high frequency alpha). An analysis of simple effects did not reveal any significant effects, however we found a trend level difference (p=0.097) between low and high-frequency alpha for the positive voxel sub-selection. This indicates a stronger negative relationship between upper alpha and the positive BOLD signal as compared to lower alpha. We interpret this result as partial evidence for a feature-related contribution of the upper alpha band. “Active” cortical patches (positive BOLD) are most likely involved in the processing of visual features (irrespective of the specific feature), and additionally a more general (possibly attention-related) activation. In turn the negative BOLD signal might contain less feature-specific activation and is most likely related to attention-driven deactivation. The negative relationship between alpha and negative BOLD could thus be interpreted as an indirect effect, resulting from a reduced alpha decrease in cortical patches responding to non-attended receptive field locations. However, the involvement of attention-related processes remains speculative, since attention was not explicitly manipulated as part of the experiment.

      We have added Figure 4 B.

      We have also added this section to the Results:

      “We furthermore found a significant interaction (p<sub>FDR</sub><0.05) between positive or negative BOLD signal change and lower or upper α sub-bands (8 - 10 or 11 - 13 Hz respectively) by means of a linear mixed effects model. An analysis of simple effects revealed that upper α frequencies are stronger negatively related to the positive BOLD signal as compared to lower α on a trend level (p<sub>FDR</sub> = 0.097).”

      And the following section of the Discussion was extended:

      “The significant interaction between the sign of the BOLD signal deflection and upper or lower α bands (see Figure 4 B) further indicates that multiple α-related processes contribute differentially to positive or negative BOLD. "Active" cortical patches (positive BOLD) are most likely involved in the processing of visual features (irrespective of the specific feature), and additionally a more general (possibly attention-related) activation. In turn the negative BOLD signal might contain less feature-specific activation and is most likely related to attention-driven deactivation. This hypothesis receives additional support from the trend-level difference in α sub-bands for positive BOLD, indicating that lower α is less related to the active, possibly feature-related processes. The absence of this difference for negative BOLD again indicates a broader, more general process. Future experiments manipulating visual features and attention might reveal a differential upper and lower α response to attended visual features and a more general relationship between α (and possibly superficial layer cortical activity) for suppressed (unattended) receptive fields.”

      (3) In the discussion (line 330 onwards), the authors mention that low-frequency alpha is predominantly related to superficial layers, referencing Figure 4A. I have a hard time appreciating this pattern there. Can the authors provide some more information on where to look?

      We thank the reviewer for pointing out the lack of clarity of this section in the Discussion. We have now rephrased the Discussion, focusing more on the laminar difference and keeping the frequency difference to a separate paragraph. Our main argument for possibly multiple alpha-related processes are twofold: a difference in alpha frequency depending on the underlying analysis (low vs high frequency alpha) and a different layer distribution (superficial layers vs. superficial and deep layers, depending on the analysis). The respective section in the Discussion now focuses on the laminar difference only. We find a negative relationship between alpha and the BOLD signal most prominently in superficial layers (feature-unspecific contrast for the BOLD signal with negative t-values; Figure 4). In addition, we find a superficial and deep layer contribution for the feature-specific contrast (congruent - incongruent; Figure 5A). While the here presented experiment was set out to investigate feature-specific processes, the meaning of the feature-unspecific results are of speculative nature. Future experiments should target the laminar difference between feature-specific and unspecific processes with respect to alpha frequency and layer distribution directly. 

      We have modified the respective sections in the Discussion:

      “Furthermore, we observed that the relationship between the feature-specific BOLD signal and α is predominantly linked to frequencies above 11 Hz (see Figure 5A). An analysis of upper and lower α sub-bands revealed a significant interaction between congruence condition and α frequency. This interaction was mainly driven by the upper α band (11 to 13 Hz). For congruently selected voxels, the negative relationship was significantly stronger (over lower α), while for incongruent selection it was significantly weaker. No such difference has been observed for the lower α component, which indicates a more feature-specific involvement of upper α and a more general modulatory effect for lower α frequencies. Since individual frequency variations have been included as a random slope in the linear mixed-effects model, these effects cannot be explained by a subset of participants driving lower or upper α separately. Specifically our findings on upper α indicate that α is not exclusively linked to global signal modulations, which has been the traditional perspective [...]”

      “Not only did we find a dissociation in the frequency domain between the relationship of α and the BOLD signal, but furthermore found that the laminar activation patterns provide further evidence for potentially multiple α-related processes. The association between α and the BOLD signal was strongest in superficial layers for negative BOLD activity and feature-specific activity (see Figure 4A and 5A). However, deep layer-related α effects were limited to feature-specific processes only (see Figure 5 A Co-Inco). These findings suggest that superficial layer α reflects are broader, more general process, while deep layer α operates more narrowly, linked to the processing of the visual features themselves. Previous findings using laminar fMRI (which did not include the investigation of oscillatory activity), indicate that superficial layer activity might be more related to the modulation of attention [...]”

      (4) How did the authors deal with the signal-to-noise ratio (SNR) across layers, where the presence of larger drain veins typically increases BOLD (and thereby SNR) in superficial layers? This may explain the pattern of feature-unspecific effects in the alpha (Figure 3). Can the authors perform some type of SNR estimate (e.g., split-half reliability of voxel activations or similar) across layers to check whether SNR plays a role in this general pattern?

      We agree with the reviewer that the vascular draining effect typically leads to increased signal change in superficial layers, the effect on (t)SNR however might be less straightforward. We did not include any counteracting measures, because we were not interested in the amplitude of the signal change, but now include an estimate of tSNR (See Figure S10 in Supplementary Figures). We found that in fact the signal-to-noise ratio is higher in deep layers. Most importantly however, the tSNR layer profiles we identified do not reflect the correlation layer result patterns of the combined EEG-fMRI analysis. This indicates that our results are most likely not the result of tSNR differences. In order to confirm our tSNR pattern we have also conducted a second layer analysis based on the LAYNII toolbox, which assigns voxels between pial and white matter to distinct layers (as compared to our fraction-based approach) and found a similar profile as with our initial analysis. However, absolute tSNR values were found to be higher for our weighted layer analysis. We speculate that while functionally relevant components of the BOLD signal drain towards superficial layers, physiological noise components will drain towards superficial layers as well.

      It is furthermore worth pointing out that for the contrast (congruent - incongruent), the vascular draining effect would cancel out between the conditions. Our findings on superficial and deep layers for those contrasts can hence not be explained by vascular draining at all.

      We have added Figure S10 to Supplementary Figures and the following section to the Discussion:

      “A major concern for laminar fMRI is the vascular draining effect (Markuerkiaga et al., 2016), which typically leads to increased signal amplitudes closer to the surface. Here, we did not investigate the signal change per se, but rather the relationship with EEG power changes. To ensure that the results do not stem from differences in tSNR across layers, we conducted a tSNR analysis (see Figure S10 in Supplementary Figures). We found that the highest tSNR was obtained from deep layers, as compared to middle and superficial layers. To verify, we computed the tSNR using a second layering algorithm (LayNii, see Huber et al. 2021), which yielded lower absolute values, but a comparable layer profile. The obtained tSNR is not reflected in any of our result profiles (see Figures 4 and 5), which strengthens the validity of the here presented results. We speculate that tSNR in deep layers is higher, because both functionally relevant components of the BOLD signal and physiological noise components drain towards superficial layers.”

      (5) The GLM used for modelling the fMRI data included lots of regressors, and the scanning was intermittent. How much data was available in the end for sensibly estimating the baseline? This was not really clear to me from the methods (or I might have missed it). This seems relevant here, as the sign of the beta estimates plays a major role in interpreting the results here.

      This is a very important remark and we would like to apologise for the confusion. It was not clear in the manuscript that the GLM was computed on z-transformed fMRI data. We have not specifically collected any “baseline volumes”. A positive beta value would indicate that the sign of the predictor matches the sign of the BOLD signal deflection (and vice versa).

      We have added or modified the following sections in Results and Methods respectively:

      “Before the GLM was computed, the fMRI data was z-transformed across time, separately for each block and voxel.”

      “A general linear model (GLM) has been computed with predictors for each TF bin separately for all voxels in V1 that later have been sub-selected according to the respective condition. Time courses for each voxel have been z-transformed before the GLM was computed for each voxel and experimental block separately. Afterwards, each of the resulting regression coefficients (β values) were multiplied with the voxel-specific layer weights that have been obtained as described above.”

      (6) Some recent research suggests that gamma activity, much in contrast to the prevailing view of the mechanism for feedforward information propagation, relates to the feedback process (e.g., Vinck et al., 2025, TiCS). This view kind of fits with the localization of gamma to the deep layer here?

      (7) Another recent review (Stecher et al., 2025, TiNS) discusses feature-specific codes in visual alpha rhythms quite a bit, and it might be worth discussing how your results align with the results reported there.

      We would like to thank the reviewer for pointing out these papers. Yes, we believe that those could be very related to the effects reported here. At the time of writing the initial manuscript we were not aware of the mentioned publications. 

      We have now included these papers in the Discussion:

      “Recent publications on the information exchange within and between primary visual cortex areas of macaques also reported deep layer γ band activity depending on the stimulus material (Gieselmann et al., 2022; Ferro et al., 2021). Those publications challenge the feed-forward exclusivity of γ altogether by revealing intra-area feedback communication in V1 from layer 5 to layer 6 and layer 6 to supra-granular layers. Possibly, the relationship between γ and deep layer BOLD we observed is also related to similar processes (Vinck et al., 2025).”

      “Similarly, in a recent opinion article, Stecher et al. (2025) promote the idea of "content-aware" α-oscillations. In agreement with our results, the authors argue that α-oscillations are related to content-specific feedback signals, reflected in increased decoding performance based on α power of top-down related processes, even prior to the onset of the stimulus (Hetenyi et al., 2025).. Accordingly, we interpret the lower α effect [...]”

      Reviewer #2 (Public review):

      The authors address a long-standing controversy regarding the functional role of neural oscillations in cortical computations and layer-specific signalling. Several studies have implicated gamma oscillations in bottom-up processing, while lower-frequency oscillations have been associated with top-down signalling. Therefore, the question the authors investigate is both timely and theoretically relevant, contributing to our understanding of feedforward and feedback communication in the brain. This paper presents a novel and complicated data acquisition technique, the application of simultaneous EEG and fMRI, to benefit from both temporal and spatial resolution. A sophisticated data analysis method was executed in order to understand the underlying neural activity during a visual oddball task. Figures are well-designed and appropriately represent the results, which seem to support the overall conclusions. However, some of the claims (particularly those regarding the contribution of gamma oscillations) feel somewhat overstated, as the results offer indeed some significant evidence, but most seem more like a suggestive trend. Nonetheless, the paper is well-written, addresses a relevant and timely research question, introduces a novel and elegant analysis approach, and presents interesting findings. Further investigation will be important to strengthen and expand upon these insights.

      One of the main strengths of the paper lies in the use of a well-established and straightforward experimental paradigm (the visual oddball task). As a result, the behavioural effects reported were largely expected and reassuring to see replicated. The acquisition technique used is very novel, and while this may introduce challenges for data analysis, the authors appear to have addressed these appropriately.

      Later findings are very interesting, and mainly in line with our current understanding of feedback and feedforward signalling. However, the layer weight calculation is lacking in the manuscript. While it is discussed in the methods, it would help to briefly explain in the results how these weights are calculated, so that the reader can better follow what is being interpreted.

      Line 104 states there is one virtual channel per hemisphere for low and high frequencies. It may be helpful to include the number of channels (n=4) in the results section, as specified in the methods. Also, this raises the question of whether a single virtual channel (i.e., voxel) provides sufficient information for reproducibility.

      We thank the reviewer for encouraging us to clarify the virtual channel selection and we agree that the current description could be misleading. Indeed, we selected 4 virtual channels in total: 1 for each frequency band (alpha/gamma), for each hemisphere separately. The main goal of this selection was to find the clearest response of that frequency band to the task. Previous publications used a supervised (ICA-based) approach to extract those responses. To increase reproducibility, we have chosen an unsupervised beamformer-based approach. The reconstruction of time or frequency-resolved sources in the brain typically yields spatially highly correlated results. Publications focusing on this type of analyses report a spatial extent of typically multiple centimetres, which here is the case as well (see Figure 3A of the updated manuscript). As such, the single voxel selection boils down to selecting the peak response within a large patch of very similarly responding voxels. Using this approach we were able to select the frequency response with the highest possible SNR. We do not however claim that the respective single voxel is exclusively carrying this information. In addition we have added a short explanation to the Discussion, since we believe that virtual channel selection with a different objective (e.g. maximising the difference between conditions or maximising cross-frequency coupling, etc.) could indeed profoundly impact the EEG-fMRI correlation, which would open up opportunities for interesting analyses that are however beyond the scope of this project.

      We have added the following section to the Discussion:

      “Future work might also vary the exact virtual channel selection for obtaining EEG-based regressors. Here, we focused on the grid points (voxel locations) with the strongest α or γ response for each frequency band in each hemisphere, derived from the average frequency response to maximise SNR. However, selecting the respective virtual channels based on the response to specific stimulus features or the interaction between high and low frequency bands are possibilities worth exploring in future work.”

      One area that would benefit from further clarification is the interpretation of gamma oscillations. The evidence for gamma involvement in the observed effects appears somewhat limited. For example, no significant gamma-related clusters were found for the feature-unspecific BOLD signal (Figure 2). Significant effects emerged only when the analysis was restricted to positively responding voxels, and even then, only for the contrast between EEG-coherent and EEG-incoherent conditions in the feature-specific BOLD response. It remains unclear how to interpret this selective emergence of gamma-related effects. Given previous literature linking gamma to feedforward processing, one might expect more robust involvement in broader, feature-unspecific contrasts. The current discussion presents the gamma-related findings with some confidence, and the manuscript would benefit from a more nuanced reflection on why these effects may not have appeared more broadly. The explanation provided in line 230, that restricting the analysis to positively responding voxels may have increased the SNR, is reasonable, but it may not fully account for the absence of gamma effects in V1's feature-unspecific response. Including the actual beta values from Figure 4 in the legend or main text would also help readers better assess the strength and specificity of the reported effects.

      We agree with the reviewer that the missing gamma-band response for the feature-unspecific signal, as well as the limitation of the effect solely to the feature-specific contrast for positive voxel selections only was unexpected. In fact, based on previous literature, we were expecting a feature-unspecific effect in the gamma band as well. However, the literature on laminar level EEG-fMRI is sparse and previous experiments used tasks that did not allow for the separation into distinct features (here left or right-oriented gratings). While we cannot fully explain the absence of the gamma effect for the feature-unspecific condition, we reasoned that our stimuli evoked weaker gamma band responses compared to previous literature. 

      The fact that we only see a significant gamma band response for the contrast for positive voxel selections can be interpreted twofold: First, previous experiments limit their analyses to positive BOLD responses only, for which we find an effect as well. Second, the fact that a significant effect could only be obtained for the contrast, might indicate that gamma band activity is related to the actual features themselves. A cortical column responding to left-oriented gratings would then be related to a gamma band response linked to that orientation. If this response to a single orientation could not be fully captured due to SNR-related issues, we would not see this effect in the congruent-only condition and also not in the feature-unspecific condition (because this boils down to both congruent conditions combined). If gamma-band oscillations are actually reflecting the response of a column to a certain orientation, then the lowest possible response would be found for the exact orthogonal orientation (here the incongruent condition). The contrast between most preferred and most not-preferred orientation might have helped to overcome the inherently low SNR, explaining the results for the contrast.

      Lastly, we did not include actual beta values in the main text, because those might be misleading. We compute the relationship between EEG power and the BOLD signal for every voxel separately, then weighted the result with the respective layer weight and lastly aggregated across voxels.This means that the beta values express the strength of the association between EEG and fMRI for an average voxel. For this reason the values are tiny and the values themselves are less meaningful than “typical” beta values.

      We have added or modified the following sections in the Discussion or Methods respectively:

      “Based on previous literature, we expected a γ band effect for the congruent condition of the feature-specific analysis (Scheeringa et al., 2016), which we did not observe. A possible explanation could be the used stimulus material in our experiment as compared to Scheeringa et al., (2016). Muthukumaraswamy et al., (2013) found that stationary gratings evoke a weaker γ band response as compared to moving annular stimuli that have been used by Scheeringa and colleagues. If γ is related to the processing of the actual features themselves (e.g. to a column preferably responding to left-oriented gratings), then contrasting congruent and incongruent voxel selections provides the largest possible contrast-to-noise ratio (CNR). In turn annular stimuli as previously used might have activated all possible orientations and thus might have greatly boosted γ SNR.”

      “The described procedure of computing a GLM based on z-transformed data using z-transformed predictors yields β-coefficients that reflect the average relationship of a single voxel's BOLD response for a given layer (fraction of the single voxel's β) with EEG power changes of a specified frequency.”

      Relating to behavioural findings for underlying neural activity, could the authors test on a trial-by-trial basis how behavioural performance relates to the BOLD signal / oscillatory activity change? Line 305 states that "Since behavioural performance in the present study was consistently high at 94% on average and participants were instructed to respond quickly to potential oddball stimuli, a higher alpha frequency might reflect a more successful stimulus encoding and hence faster and more accurate behavioural performance." Also, this might help to relate the findings to the lower vs upper alpha functionality difference.

      This is a very interesting suggestion. We now include an exploratory analysis of the relationship between frequency and behavioural performance in the Supplementary Figures (see Figure S12). We did not perform a correlation between behavioural performance and alpha over trials because of the low numbers of oddball trials (N=40) and very limited number of false responses (94% response accuracy on average). However, we computed a correlation across participants. After averaging the alpha time-frequency spectrum across non-oddball trials, the individual alpha frequency was determined by the frequency where the alpha decrease (between 0.1 and 0.8 s post-stimulus) was largest. The correlation between alpha frequency and either reaction times and d’ (as a measure for accuracy), yields a significantly positive relationship between d’ and alpha frequency. This indicates that alpha frequency is related to task performance. We interpret those exploratory findings such that high behavioural accuracy is reflected by a stronger modulation of high-frequency alpha power. 

      We have added Figure S12 to Supplementary Figures.

      We have also added the following sections to Results and Discussion respectively:

      “An exploratory analysis of the relationship between individual α frequency (IAF) and task performances underlines this finding (see Figure S12 in Supplementary Figures). Thereby the IAF was obtained from the average α power spectrum of each participant. The frequency with the strongest decrease between 0.1 and 0.8 s after stimulus onset served as the IAF. We correlated IAF with average response times to correct oddball trials and d' as a measure for accuracy and found a significant positive correlation between IAF and d' (p < 0.05).”

      “We exploratively correlated the average IAF during non-oddball trials with the average task accuracy (d') across participants and indeed found IAF and task performance to be positively correlated (See Figure S12 in Supplementary Figures).”

      In Figure 4, the EEG alpha specificity plot shows relatively large error bars, and there is visible overlap between the lower and upper alpha in both congruent and incongruent conditions. While upper alpha shows a positive slope across conditions and lower alpha remains flat, the interaction appears to be driven by the change from congruent to incongruent in upper alpha. It is worth clarifying whether the simple effects (e.g., lower vs upper within each condition) were tested, given the visual similarity at the incongruent condition. Overall, the significant interaction (p < 0.001, FDR-corrected) is consistent with diverging trends, but a breakdown of simple effects would help interpret the result more clearly. Was there a significant difference between lower and upper alpha in congruent or incongruent conditions?

      We thank the reviewer for this important remark and have added a simple effects analysis (see Figures 4 b and 5 b, e). We found that the main driver for the interaction between congruence condition and alpha frequency is upper alpha. Specifically the negative relationship between upper alpha and the BOLD signal is significantly stronger for the congruent condition and weaker for the incongruent condition. This indicates the upper alpha indeed is related to the processing of visual features.

      We have added a simple effects analysis (See Figures 4 and 5).

      We have added or modified the following in Results, Discussion and Methods respectively:

      In Results:

      “We furthermore found a significant interaction (p<sub>FDR</sub> < 0.05) between positive or negative BOLD signal change and lower or upper α sub-bands (8 - 10 or 11 - 13 Hz respectively) by means of a linear mixed effects model. An analysis of simple effects revealed that upper α frequencies are stronger negatively related to the positive BOLD signal as compared to lower α on a trend level (p<sub>FDR</sub> = 0.097).”

      “After correcting for multiple comparisons, we found a significant interaction (p<sub>FDR</sub> < 0.001). This interaction is mainly driven by the upper α sub-band, as indicated by the simple effects analysis. We found a significantly stronger negative relationship of upper α and the BOLD signal for congruent selections (p<sub>FDR</sub> < 0.01) and the reverse for the incongruent condition (p<sub>FDR</sub> < 0.01), as well as a significantly stronger negative relationship within the upper α sub-band for congruent over incongruent voxel selections (p<sub>FDR</sub> < 0.01).”

      “This interaction is mainly driven by the upper α sub-band, as indicated by the simple effects analysis, which revealed a significantly stronger negative relationship of upper α and the BOLD signal for congruent over incongruent selections (p<sub>FDR</sub> < 0.001).”

      In Discussion:

      “An analysis of upper and lower α sub-bands revealed a significant interaction between congruence condition and α frequency. This interaction was mainly driven by the upper α band (11 to 13 Hz). For congruently selected voxels, the negative relationship was significantly stronger (over lower α), while for incongruent selection it was significantly weaker. No such difference has been observed for the lower α component, which indicates a more feature-specific involvement of upper α and a more general modulatory effect for lower α frequencies.”

      In Methods:

      “Significant interactions were decomposed into simple effects using Wald tests on the model coefficients, ensuring that post-hoc comparisons were derived from the same statistical global variance as the primary interaction.”

      Overall, this study provides a valuable contribution to the literature on oscillatory dynamics and laminar fMRI, though some interpretations would benefit from further clarification or qualification.

      Reviewer #3 (Public review):

      Summary:

      Clausner et al. investigate the relationship between cortical oscillations in the alpha and gamma bands and the feature-specific and feature-unspecific BOLD signals across cortical layers. Using a well-designed stimulus and GLM, they show a method by which different BOLD signals can be differentiated and investigated alongside multiple cortical oscillatory frequencies. In addition to the previously reported positive relationship between gamma and BOLD signals in superficial layers, they show a relationship between gamma and feature-specific BOLD in the deeper layers. Alpha-band power is shown to have a negative relationship with the negative BOLD response for both feature-specific and feature-unspecific contrasts. When separated into lower (8-10Hz) and upper (11-13Hz) alpha oscillations, they show that higher frequency alpha showed a significantly stronger negative relationship with congruency, and can therefore be interpreted as more feature-specific than lower frequency alpha.

      Strengths:

      The use of interleaved EEG-fMRI has provided a rich dataset that can be used to evaluate the relationship of cortical layer BOLD signals with multiple EEG frequencies. The EEG data were of sufficient quality to see the modulation of both alpha-band and gamma-band oscillations in the group mean VE-channel TFS. The good EEG data quality is backed up with a highly technical analysis pipeline that ultimately enables the interpretation of the cortical layer relationship of the BOLD signal with a range of frequencies in the alpha and gamma bands. The stimulus design allowed for the generation of multiple contrasts for the BOLD signal and the alpha/gamma oscillations in the GLM analysis. Feature-specific and unspecific BOLD contrasts are used with congruently or incongruently selected EEG power regressors to delineate between local and global alpha modulations. A transparent approach is used for the selection of voxels contributing to the final layer profiles, for which statistical analysis is comprehensive but uses an alternative statistical test, which I have not seen in previous layer-fMRI literature.

      A significant negative relationship between alpha-band power and the BOLD signal was seen in congruently (EEGco) selected voxels (predominantly in superficial layers) and in feature-contrast (EEGco-inco) selected (superficial and deep layers). When separated into lower (8-10Hz) and upper (11-13Hz) alpha oscillations, they show that higher frequency alpha showed a significantly stronger negative relationship with congruency than lower frequency alpha. This is interpreted as a frequency dissociation in the alpha-BOLD relationship, with upper frequency alpha being feature-specific and lower frequency alpha corresponding to general modulation. These results are a valuable addition to the current literature and improve our current understanding of the role of cortical alpha oscillations.

      There is not much work in the literature on the relationship between alpha power and the negative BOLD response (NBR), so the data provided here are particularly valuable. The negative relationship between the NBR and alpha power shown here suggests that there is a reduction in alpha power, linked to locally reduced BOLD activity, which is in line with the previously hypothesized inhibitory nature of alpha.

      Weaknesses:

      It is not entirely clear how the draining vein effect seen in GE-BOLD layer-fMRI data has been accounted for in the analysis. For the contrast of congruent-incongruent, it is assumed that the underlying draining effect will be the same for both conditions, and so should be cancelled out. However, for the other contrasts, it is unclear how the final layer profiles aren't confounded by the bias in BOLD signal towards the superficial layers. Many of the profiles in Figure 3 and Figure 4A show an increased negative correlation between alpha power and the BOLD signal towards the superficial layers.

      We thank the reviewer for this important remark. Reviewer 1 raised a similar concern and I would like to refer you to our response to Reviewer 1, point 4. The veinal draining typically results in a higher signal change closer to the cortical surface. We did not take any measures to counteract this effect, but provide an analysis of tSNR in Supplementary Figures (see Figure S10). Possibly due to the drainage of physiological noise towards the surface, we found the highest tSNR in deep, followed by middle and superficial layers. To verify those results we computed the same analysis using a second layering algorithm, which resulted in the same profile, but overall less tSNR. Crucially the tSNR profile is not reflected in our EEG-fMRI results.

      We have added Figure S10 to Supplementary Figures and the following section to the Discussion:

      “A major concern for laminar fMRI is the vascular draining effect (Markuerkiaga et al., 2016), which typically leads to increased signal amplitudes closer to the surface. Here, we did not investigate the signal change per se, but rather the relationship with EEG power changes. To ensure that the results do not stem from differences in tSNR across layers, we conducted a tSNR analysis (see Figure S10 in Supplementary Figures). We found that the highest tSNR was obtained from deep layers, as compared to middle and superficial layers. To verify, we computed the tSNR using a second layering algorithm (LayNii, see Huber et al. 2021), which yielded lower absolute values, but a comparable layer profile. The obtained tSNR is not reflected in any of our result profiles (see Figures 4 and 5), which strengthens the validity of the here presented results. We speculate that tSNR in deep layers is higher, because both functionally relevant components of the BOLD signal and physiological noise components drain towards superficial layers.”

      When investigating if high alpha (8-10 Hz) and low alpha (11-13 Hz) are two different sources of alpha, it would be beneficial to show if this effect is only seen at the group level or can be seen in any single subjects. Inter-subject variability in peak alpha power could result in some subjects having a single low alpha peak and some a single high alpha peak rather than two peaks from different sources.

      We agree with the reviewer that a bias in a subset of participants to generally higher or lower alpha frequencies could potentially skew the presented results. While the initially computed model included a random intercept for the frequencies, we have now added the random slope as well. This ensures that the difference between low and high frequency alpha is indeed only driven by the difference in condition and not the result of individual differences across conditions themselves.

      In order to verify that not a small subset of participants is driving the result pattern, we also computed the fraction of participants that either show the dual alpha pattern (i.e. follow the exact pattern of the group average), contribute to the group average with a single peak or contradict the pattern entirely. Thereby, 40.4% of all participants show a dual alpha pattern, 38.4% a single alpha pattern in the direction of the group average and 21.2% contradict the group average. See Author response image 1:

      Author response image 1.

      Alpha Response Patterns with Example Subjects: V1 Feature Specific Contrast

      We would also like to highlight our added exploratory analysis of the relationship between alpha frequency and behavioural performance, which was requested by Reviewer 2, point 3. We find a significant positive correlation between alpha frequency and task performance on a group level. This indicates that higher alpha frequencies might be related to better discrimination of visual features. We speculate that participants with better task performance are capable of modulating their upper alpha more than participants with worse performance.

      We have added Figure S12 to Supplementary Figures.

      We have also added the following sections to Results and Discussion respectively:

      “An exploratory analysis of the relationship between individual α frequency (IAF) and task performances underlines this finding (see Figure S12 in Supplementary Figures). Thereby the IAF was obtained from the average α power spectrum of each participant. The frequency with the strongest decrease between 0.1 and 0.8 s after stimulus onset served as the IAF. We correlated IAF with average response times to correct oddball trials and d' as a measure for accuracy and found a significant positive correlation between IAF and d' (p < 0.05).”

      “We exploratively correlated the average IAF during non-oddball trials with the average task accuracy (d') across participants and indeed found IAF and task performance to be positively correlated (See Figure S12 in Supplementary Figures).”

      The figure layout used to present the main findings throughout is an innovative way to present so much information, but it is difficult to decipher the main findings described in the text. The readability would be improved if the example (Appendix 0 - Figure 1) in the supplementary material is included as a second panel inside Figure 3, or, if this is not possible, the example (Appendix 0 - Figure 1) should be clearly referred to in the figure caption. 

      Since Reviewer 1 suggested using an entirely different figure layout, we now opted to remove some information from the main text figures (we only show the 10% threshold, but 5% and 25% is in Supplementary Figures) and chose a more common figure layout. See Figures 4 and 5.

      Recommendations for authors:

      Reviewer #2 (Recommendations for the authors):

      The contrasts used in the analysis are not clearly introduced in the main text. While the methods section explains them more thoroughly, some of this explanation would be better placed in the results section, where the contrasts are first used. Specifically, the concepts of "feature-specific" vs. "feature-unspecific" BOLD signals are introduced with a very brief definition, which could be confusing for readers. The same applies to the terms EEG co and EEG inco; it would help to briefly explain these when they are first mentioned in the results. The supplementary figures and legends are helpful, so it is clear that the authors were prioritising clarity overall.

      The respective analyses are now also explained in the Results section:

      “During each trial either a left or a right-oriented grating was presented, from which two types of analyses have been derived: feature-unspecific BOLD activation (i.e. the response to any stimulus orientation), and feature-specific BOLD activation (i.e. the response to a specific stimulus orientation or the contrast between them). Thereby, fMRI data and EEG-based regressors could either be combined congruently (Co) by combining the BOLD signal of orientation-selective voxels with EEG-based regressors built from the same orientation trials, or incongruently (Inco), by combining the orientation-specific BOLD signal with EEG-based regressors built from the other orientation trials. Finally, those two congruency conditions have been contrasted (Co-Inco).”

      Figures are overall clear and illustrative of the results. For Figure 4, however, the use of dotted elements makes it somewhat harder to interpret what's being shown. While the supplementary figure clarifies the findings, rephrasing the figure legend to explain what the dotted lines represent would be helpful.

      Figures 4 and 5 have been replaced with a new layout and legends have been improved.

      The reported ranges overlap (e.g., alpha: 2-32 Hz; gamma: 20-120 Hz). It would be helpful to explain why such overlapping bands were chosen.

      Both frequency bands of interest differ slightly in their later time-frequency analysis (i.e. number of tapers and filter type). The overlap itself is not meaningful per se and results from the selection of a wide band for each respective sub-band. This wide selection was chosen to avoid filter artefacts. For the alpha sub-band, we also wanted to ensure that the beta spectrum is covered which also includes the alpha harmonic and for the gamma band that the full range of high-frequency activity is captured (e.g. EMG activity).

      Only a single time point was used for baseline correction of the low alpha band. Is this typical? The authors note that due to the gradient artefact arising in the pre-stimulus period, the baseline correction is somewhat difficult, although further clarification would be useful here.

      Relatedly, was pilot scanning conducted? If so, was the presence of strong gradient artefacts unexpected? More details about this would strengthen the methodological transparency.

      Indeed only a single time bin was used as the baseline for the alpha sub-band. After the piloting phase a slight adjustment to the final fMRI sequence has been made which was not expected to introduce gradient artefacts so close to the onset of the stimulus. Unexpectedly, those artefacts were visible until 300 ms before the onset of the stimulus. Similarly, a pre-stimulus alpha was observed (starting 250 ms before the onset of the stimulus), which we also aimed to exclude from the baseline period. In the end only the time bin centered at 300 ms prior to stimulus onset was chosen. However, this time bin contains 400 ms of data (the width of the window for the time frequency analysis). Thus, the term time point was misleading, because the actual time window that made up the baseline is 500 ms to 100 ms prior to the onset of the stimulus. 

      We have adjusted our wording in Methods to make this more clear:

      “For this reason, the low frequency baseline period comprised only a single 400 ms time bin centred around -0.3 s, because a pre-stimulus α decrease was expected starting around 0.25 s prior to stimulus onset.”

      Including a one-sentence explanation of the AROS test in the main text for clarity. As line 796 in the methods: "Each significant cluster has been further processed by means of an auto-regressive rank order similarity (aros) test (Clausner and Gentili, 2022). The fundamental idea behind the AROS test is whether group averages (i.e. averages of the signal of cortical layer in the present case), can be ranked such that the rank order is explained significantly better by the data than it would if the average data could not be meaningfully sorted (i.e. is shuffled)."

      An explanation has been added to the Results section:

      “Each significant cluster was then averaged along the frequency dimension at the widest point to enable an auto-regressive rank order similarity (aros) test Clausner & Gentili (2022), testing the laminar activation profile. The aros test transforms the layer averages into a rank order and tests - using a permutation procedure - if the rank order of the layer averages explains the data better than a random rank order (shuffled layer labels) would.”

      Line 223: "In fact, an analysis of the relationship between the EEG signal and the BOLD signal that focused on the feature contrast only (L - R; independent of the comparison to baseline) revealed a trend-level result with an even stronger deep layer contribution as compared to superficial layers." Could you point to which figure represents this finding - Figure 4B?

      This refers to Figure 4A in the old manuscript, for the 25% threshold for the gamma band. Since now the new figures do not include the 25% threshold anymore, it refers to Figure S4i.

      The number of participants is missing from the main text. Including this in the results section would improve clarity.

      The description of our sample has been moved from Methods to Results.

      Given the complexity of the data acquisition and analysis, the well-designed and easy-to-follow analysis pipeline figure (currently in the supplement) would be better placed in the main text.

      The mentioned Figure has been moved to the main text (now Figure 2).

      Also, simply out of curiosity, what do the authors think about the theta blob around 200ms post-stimulus?

      The theta blob most likely reflects the post-stimulus ERP as often observed in response to visual stimuli. We hypothesise that it is stronger in the middle and superficial layers, but we did not want to extend too much the scope of this paper. Additional analyses could be performed in the future on this evoked activity.

      Reviewer #3 (Recommendations for the authors):

      (1) Minor Corrections to the text and figures:

      We would like to thank the reviewer for the very valuable recommendations. Below we shortly describe how each suggestion has been implemented.

      We have made the white box more clear (see Figure 3 B).

      (b) Page 10: Top of 2nd paragraph - 'The full experimental protocol comprised a high resolution anatomical T1 scan lasting for 8 min'. The methods state this scan is 6 min 31 sec.

      The confusion results from the fact that the T1 scan was recorded during a short practice block that the participants performed inside the scanner. This block lasted 8min during which the 6 min 31 sec T1 scan was recorded. We have made this more clear:

      “Once prepared, the participant was placed inside the scanner and performed an 8 min practice block. A T1-weighted scan was acquired during this time in the sagittal orientation using a 3D MPRAGE sequence Brant-Zawadzki et al., (1992) with the following parameters: TR/TI = 2.2/1.1 s, 11° flip angle, FOV 256 x 256 x 180 mm and an 0.8 mm isotropic resolution. Parallel imaging (iPAT = 2) was used to accelerate the acquisition, resulting in an acquisition time of 6 min and 31s.”

      (c) Page 10: 'Stimulus presentation' paragraph - 'Stimuli were projected onto a screen behind the subject's head using'. The use of 'subject' should be replaced with 'participant' throughout.

      We have corrected the phrasing.

      (d) Page 14: Figures 2A and 2B are referred to incorrectly as being in the supplementary material.

      We have corrected the mistake.

      (e) Figure 5 caption: 'Regressors are build for each time-frequency bin separately.' should be 'built'

      We have corrected the mistake.

      (f) Page 16, final paragraph: 'Afterwards, each of the resulting regression coefficients (B coefficients) was multiplied with the voxel specific layer weights that have been obtained as described above.' Should be 'were multiplied'

      We have corrected the mistake.

      (g) Page 17: 'Subsequently, separate analyses were done for two frequency of interest (FOI) ranges centerd around' - typo

      We have corrected the mistake.

      (h) Page 17 - 'Within these frequency ranges inferential statistics based a cluster level' - missing word. Should be 'based on a cluster level'

      We have corrected the mistake.

      (i) Page 14 Figure 2B and 5D are referred to incorrectly as being in the supplementary material.

      We have corrected the mistake.

      (2) fMRI data pre-processing:

      Please provide a comment on the EEG-fMRI data quality - e.g. tSNR of EPI data. Perhaps example EPI data could be shown in the supplementary information.

      We included the below Figure S11 in Supplementary Figures showing an example EPI. We have also included an illustration of the result of our layering approach. Furthermore, we included a tSNR analysis (see Response to Reviewer 1, point 4).

      On a practical note - with 14-minute long runs whilst wearing an EEG cap, I would expect participant motion to be a concern. Could you provide some metrics on perhaps the average of the mean and maximum per subject displacement/rotation?

      We ensured that participants receive tactile feedback for their respective head motion from a strip of tape span across their foreheads. This resulted in overall manageable motion during each experimental block. During the main experiment, the average framewise displacement was 0.3 mm, with an average total translation of 1.6 mm and an average total rotation of 1.6 deg within each block. 

      We have added Figure S13 to Supplementary Figures.

      We have added a section to Methods:

      “Subject motion per block was low, with a mean (SD) frame-wise displacement Power et al. (2012) of 0.34 mm (0.24 mm) for the main experiment and 0.23 mm (0.22 mm) for the retinotopy (see also Figure S13 in Supplementary Figures).”

    1. eLife Assessment

      The report by Liu and colleagues provides a valuable analysis of environmental adaptation across diverse lineages of the grass Phragmites australis differing by their level of ploidy. The analysis reports solid evidence that lineages with distinct levels of ploidy occupy different climate niches. The use in tandem of regional survey and common garden experiment represents a convincing approach to suggest a correlation between ploidy and climate adaptation. This manuscript will be of interest to a broad community of ecological genomicists interested in how structural variation in gene dosage potentially affects the pattern of adaptation.

    2. Reviewer #1 (Public review):

      Summary:

      The article is testing the relative advantages of plant lineages with differing ploidy and admixture across environmental gradients. The results show that intraspecific variation in ploidy and admixture between lineages impacts plant traits that may enable persistence and range expansion.

      Strengths:

      Suitable marker panel size and convincing results that include attempts to analyse mixed ploidy level data, which is a challenge.

      Weaknesses:

      (1) Inadequate explanation of allele dosage for ploidy levels, some of which do not match the allele counts expected for genome copy number.

      (2) The setup and sample sizes of the common garden experiments are very unclear. The numbers implied are extremely low to draw robust conclusions.

      (3) Unclear how allele dosage is determined. Given it's so central to many analyses, it would be useful to see how this is done rather than use a citation.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript describes a combination of species distribution mapping experimental data from common garden and physiological experiments to project the future distribution of genetic subgroups with the widespread grass Phragmites australis. Overall, the sample sizes seem appropriate for the questions being asked, and the key results regarding projected change in distribution of the focal lineages are well supported. However, at this point, it is difficult to evaluate the broader impact of the work on the field or the utility of the data for the broader community outside of those studying the focal species, P. austrina.

      Strengths:

      A key strength of the paper is the use of common garden and physiological experiments in conjunction with species distribution modeling. The experiments provide a mechanistic basis for the correlations between interspecific lineage and climatic data, suggesting that the distributional patterns are more likely to result from genetic differences rather than limited dispersal among regions. I would, in fact, emphasize the experimental validation of modeling efforts even more in the introduction.

      Weaknesses:

      I see two weaknesses with the framing of the ms and the presentation of the results. First, no data support the claims that polyploidy has any causal effect. The ploidy levels are, in fact, completely confounded with other genetic differences, so it is not possible to eliminate genetic variation, independent of ploidy, as the causative factor. As the authors note, ploidy was not manipulated in the reported experiments. Thus, the focus on polyploidy in the introduction and elsewhere distracts from the novel and informative experiments that were conducted. Second, the manuscript indicates that intraspecific variation is critical for the evolutionary potential of a species to respond to environmental change, but intraspecific variation is seldom considered in species distribution models. To me, an assessment of evolutionary potential requires estimates of heritable genetic variation and responses to selection. The sample sizes presented here are modest to estimate heritabilities, but the manuscript could be framed with this perspective in mind. However, instead, the manuscript performs species distribution modeling on a small number of sub-specific lineages, essentially treating them as homogeneous "species" - thus the analysis commits the same oversimplification that the manuscript highlights, but does so at a finer evolutionary scale than species. Not acknowledging this simplification (or better, examining phenotypic variation within the genetically defined lineages) hinders what would otherwise be a strength of the manuscript.

      The title suggests that asymmetric introgression and thermal tolerance are the most important findings of the work. However, the introduction contains no explanation of the potential importance of gene flow (other than to say that asymmetric gene flow was suggested by some preliminary analyses), and the discussion offers only a limited explanation of either the potential mechanisms underlying the asymmetric gene flow or its importance for the long-term evolution of the species. Similarly, the novelty of combining experiments and species distribution modeling is scarcely mentioned, and there is no exploration of the connection between tolerance alleles and gene flow. Could introgression of heat tolerance alleles alter the spread of the hybridizing lineages, for example? A greater emphasis on these general population genetic parameters could potentially highlight the broader impact of this work.

    4. Author response:

      We sincerely thank the editors and reviewers for the positive assessment of our work and for the constructive and insightful feedback.

      We fully agree with the major points raised in the public reviews and outline below our planned revisions to address them.

      Reviewer #1 raised two important concerns regarding our methodology. First, the determination of allele dosage is insufficiently explained, which is central to our ploidy assignment and downstream analyses. Second, the setup and sample sizes of the common garden experiments are unclear, raising questions about the robustness of our conclusions. We accept these criticisms and will address them as follows.

      Regarding allele dosage, we will add a detailed step-by-step description of our calling pipeline in the Methods section, including the criteria for peak height ratios and thresholds used to assign copy numbers. We will also clarify a crucial biological detail: the common reed (Phragmites australis) is an allotetraploid in its origin. As a consequence, many molecular markers, including the widely used SSR markers in previous studies, behave as disomic markers (i.e., two homeologous copies inherited in a Mendelian manner). Therefore, observing more than two alleles at a locus is indeed indicative of higher-level ploidy (hexaploidy or octoploidy) in this system. We will explicitly state this to resolve any confusion about why tetraploids in our dataset are treated as having a maximum of two alleles, while hexaploids and octoploids can carry more.

      Regarding the common garden experiment, we will explicitly report the replication number for each lineage-by-treatment combination and clarify the experimental design. We will also discuss the statistical approaches used given the sample sizes, while acknowledging that the consistency between experimental results and distributional patterns lends additional support to our conclusions.

      Reviewer #2 raised three substantive framing issues. First, ploidy is completely confounded with genetic background, yet our manuscript places undue emphasis on polyploidy as a causal factor. Second, our species distribution models treat each lineage as a homogeneous entity, failing to capture within-lineage variation and thus repeating the oversimplification we criticize. Third, we insufficiently explore the evolutionary significance of asymmetric introgression, gene flow, and the novelty of combining SDM with experiments. We fully agree with these points and will revise accordingly.

      To address the confounding issue, we will substantially reframe the manuscript to de-emphasize claims about polyploidy as a causal driver, and instead focus on the adaptive differentiation among distinct genetic lineages that happen to differ in ploidy. The Discussion will explicitly state that dissecting ploidy effects from background genetic effects will require future experimental approaches.

      To address the simplification in SDMs, we will add a clear acknowledgment of this limitation, discussing how it may affect predictive accuracy and suggesting that future studies incorporating population-level genomic data could more directly assess evolutionary potential.

      To address the insufficient exploration of introgression and the novelty of our approach, we will expand the Introduction to better highlight the value of coupling controlled experiments with SDMs at the intraspecific level. In the Discussion, we will elaborate on the evolutionary significance of asymmetric introgression, including testable hypotheses about how gene flow might mediate the spread of heat-tolerance alleles and influence lineage geographical limits under climate change.<br /> We also thank the reviewer for the suggestion to emphasize the experimental validation of SDM efforts, which we will incorporate into a revised Introduction.

      Looking beyond the present study, we envision three complementary directions that build upon our current findings. Expanding common garden experiments to include admixed individuals would test whether introgressed genomic blocks confer fitness advantages under thermal stress. Leveraging the population genomic framework established here, we will transition to whole-genome resequencing for selection scans and genotype-environment association analyses to pinpoint adaptive loci and reveal whether heat-tolerance alleles are preferentially transferred via asymmetric introgression. We will also integrate transcriptomic profiling with phenotypic measurements to identify candidate genes whose expression correlates with thermal performance and introgressed ancestry, helping to disentangle ploidy effects from genetic background. Together, these directions span expanded phenotyping, whole-genome resequencing, and transcriptome-guided discovery, forming an integrated framework that moves from the correlative patterns reported here toward mechanistic understanding. These perspectives are briefly outlined in our Discussion, and we hope the present study will serve as a foundation for these future investigations, which we plan to pursue in subsequent work.

      We believe these revisions will substantially strengthen the manuscript.

    1. eLife Assessment

      This work presents a software and hardware suite for targeted photostimulation that can be used in vivo. The package is a well-designed and documented hardware/software suite with a comprehensive build guide. This tool will likely promote important neuroscience advances through targeted real-time perturbation of the cerebral cortex. Overall, this manuscript makes a compelling case on how to design and make available power tools for the research community.

    2. Reviewer #1 (Public review):

      Lohse et al. describe an open-source system for laser scanning photostimulation (LSPS) in head-fixed animals. Although similar systems have been developed and used by different groups, Zapit provides an open-source solution requiring few custom parts and minimal coding. This tool can clearly facilitate and speed the adoption of LSPS, particularly for the increasingly used purpose of mapping the effects of focal cortical silencing during behavior. Other potential uses include mapping optogenetically evoked movements and selectively activating genetically labeled neuronal subtypes of interest in the cortex. The design is well thought through, and the presentation is mostly clear and well written.

      In general, the more modular such a system is, the better, in terms of compatibility with existing hardware and software that potential users may already have purchased - laser, galvo, and camera in particular. The system has struck a reasonable balance between allowing modularity and providing an integrated complete package, but even more flexibility would be welcome for potential users looking to cut costs, as would clearer presentation of such flexibility as already exists.

      Comments and suggestions are mostly minor, as follows.

      (1) Command signals:

      How is the relationship between analog voltage commands and laser power determined? Is this assumed (or required) to be linear (as Figure 7F implies)? Usability and modularity would be improved by an option to measure or provide a calibration curve for systems with a nonlinear mapping between command voltage and laser power.

      For the grid calibration step, how is the initial mapping from galvo voltage commands to image position determined? Presumably, some sort of initial guess or calculation based on the hardware specifications is needed for the grid calibration to be feasible. Also, how are the number of grid lines and the distance between them determined?

      Why is the mapping between analog outputs and hardware (galvos, laser, masking light) fixed? This would be trivial to make configurable and allow labs with existing setups to adopt Zapit without rewiring existing hardware.

      (2) Laser and optics:

      In Figure 1, the authors should consider explaining the scanning principle schematically, i.e., depicting how tilting of the scan mirrors translates via the scan lens into beam displacement in the specimen plane. Perhaps Zemax can be used for accurate rendering.

      Since the unexpanded beam greatly under-fills the back aperture of the lens, the z resolution is presumably terrible - which is good! That is, for the purposes of LSPS, this advantageously avoids focus-dependent effects, which might otherwise arise due to (e.g.) skull curvature. The authors should consider pointing this out, as well as providing an estimate of the z resolution.

      What is the working distance?

    3. Reviewer #2 (Public review):

      Summary:

      In this work, Lohse and colleagues develop a system for doing targeted photostimulation in mouse cortex. The system uses a camera image to target laser stimulation to stereotactically defined locations in mouse dorsal cortex.

      Strengths:

      The hardware is well designed, and the software is well documented and supported. The build guide and well-documented software package should allow for simple implementation of the technology. Without a doubt, this is a valuable community resource for the circuit neuroscience field.

      Weaknesses:

      No weaknesses were identified by this reviewer.

    1. eLife Assessment

      This important study combines peptide engineering, molecular docking, and functional assays to define the molecular basis of ligand recognition and activation of the human Y4 receptor and to identify three novel small-molecule agonists. The evidence supporting the conclusions is convincing, with complementary experimental and computational approaches providing strong support for the proposed receptor-ligand interactions. While concentration-response analyses of the small-molecule agonists and additional structural or mutagenesis studies would further strengthen the work, these are not essential to support the main conclusions. The work will be of interest to researchers studying GPCR pharmacology, structural biology, and ligand discovery.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript describes an investigation of peptide analogue agonists selective for the human Y4 receptor for pancreatic polypeptide over Y1, Y2 and Y5 receptors. After studies of mutated Y4R in transiently transfected COS-7 cells, binding models were calculated. Then, screening of a virtual library identified three non-peptidergic (albeit somewhat peptide-like) compounds with potential agonist activity that were subsequently confirmed and furthermore were found to have receptor interactions similar to the peptide analogues. This study provides fundamental new information that improves understanding of the Y4R structure and mechanism of activation by the native agonist and the selective peptide analogues. The non-peptide agonists have potential for future pharmacotherapy.

      Strengths:

      All of the experiments seem to be well performed, using state-of-the-art methods. The manuscript is quite comprehensive and has used a broad range of methods. The conclusions are convincingly supported by the experimental results.

      Weaknesses:

      The mutagenesis was almost exclusively based on the replacement of potentially interesting amino acid residues with alanine. Replacement with other residues, based on modelling and docking, could have refined the model further. Neither molecular dynamics nor cryo-EM was used to study the agonists' interactions with the Y4 receptor and these are therefore likely next steps in the characterization of the Y4R mechanism of activation.

    3. Reviewer #2 (Public review):

      Summary:

      Pelczyk et al. investigated the binding site of the neuropeptide Y Y4 receptor with the aim of identifying novel small-molecule agonists. The authors first assessed small cyclic peptides as tool compounds and then identified interactions between peptides and receptor residues, which were confirmed by single-point mutagenesis combined with functional assays for intracellular signalling. It is interesting that a peptide receptor can be activated by the relatively small cyclic peptides used in the study. The authors identified both common and peptide-specific interactions. The identified interactions guided ultra-large library screening, which yielded 53 compounds, 3 of which were confirmed as Y4R-specific agonists in an IP-one accumulation assay.

      Strengths:

      The combination of techniques (docking, mutagenesis and functional assays) strongly supports the identification and evaluation of small molecules as agonists at the neuropeptide Y Y4 receptor. Functional assays highlight residues that are important for the binding of all tested peptides, as well as residues with peptide-specific importance.

      The structure-activity relationship component of the study nicely highlights which components of the peptide are important for binding to the different members of the neuropeptide Y receptor family.

      Weaknesses:

      It would have been great to see concentration-response curves for the three identified small-molecule agonists, as this would have stengthened the case for these agonists.

    1. eLife Assessment

      This study presents a valuable metagenomic analysis of the gut microbiome in sickle cell disease (SCD) patients, revealing associations between bacteriophage, host immunity, and SCD pathophysiology. While these data are interesting and helpful for hypothesis generation, they are deemed incomplete; additional experiments would be needed to test causality and to provide mechanistic insight. Despite these limitations, this work will be of broad interest to researchers studying SCD, immunology, phage biology, and the microbiome, adding to the small but growing literature suggesting a microbial component to SCD.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, Flamholz and colleagues use metagenomic sequencing to profile the microbiome of individuals with sickle cell disease (SCD), the most common genetic blood disorder in the world. To build on previous studies that found dysbiosis in SCD, this manuscript aims to examine whether changes in either bacterial species or bacteriophages correlate with inflammatory hallmarks of the disease. The authors claim that sickle cell dysbiosis does not correlate with inflammatory hallmarks of the disease, but instead, aged neutrophil numbers and bacteriophages do. Appropriate control subjects and additional analyses are needed to support that conclusion.

      Strengths:

      The primary strength of this paper is the investigation into disease-associated changes in bacteriophages. This is an entirely novel idea in the sickle cell field, and based on the current results, may be an important, under-recognized disease hallmark. It is unclear, however, if phages are "the chicken or the egg" in terms of sickle cell inflammatory profiles; do these increases in phage number simply result from other disease processes, or are they in any way contributing to disease pathophysiology?

      Weaknesses:

      A primary weakness of the manuscript is the fact that the majority of individuals included in the control group maintain sickle cell trait (HbAS genotype). Although typically asymptomatic, it is unclear if this genotype is associated with microbial changes that would not be observed in a true control group (HbAA genotype). This is a significant limitation that may limit the ability to draw conclusions from the current data set.

      Another key weakness is the lack of beta diversity assessment. Although decreased alpha diversity is observed in individuals with SCD, and specific bacterial taxa are differentially abundant following multivariate analyses, there is no overall comparison of bacterial community composition between individuals with SCD and controls. Prior to drawing conclusions about the relationship (or lack thereof) between the SCD microbiome and inflammatory markers, it is important to know if this study did indeed find disease-associated changes in microbiome composition.

      It is unclear which individuals were used for aged neutrophil (AN) and molecular data assessments. For example, were children who were still receiving penicillin prophylaxis included in these specific assessments? Given the authors' previous work demonstrating that antibiotic treatment decreases AN pathology, it seems critical to limit all AN/molecular analyses to older subjects who are not on daily penicillin treatment (if possible).

      A minor weakness is the continued use of "disease" vs. "healthy" indicators as primary microbiome metrics that are used for molecular correlations. The lack of metric specificity - and lack of discussion regarding which diseases were used to generate these indicators (how similar/different are they to sickle cell?) - could be said to make these metrics essentially meaningless.

    3. Reviewer #2 (Public review):

      Summary:

      The study analyzes stool metagenomes from 98 SCD patients and 46 controls, with SCD and control groups matched on age, race, sex, and ethnicity. The authors report lower Shannon diversity, lower Firmicutes/Bacteroidetes ratio, loss of health-associated taxa, increased disease-associated indicators, altered butyrate/fatty-acid metabolism pathways, and enrichment of provirus/prophage fractions in SCD. They further correlate aged-like neutrophils and prophage fractions with inflammatory cytokines. The main strength is that this is not just another 16S comparison. The use of whole-community metagenomics, immune profiling, neutrophil assays, and clinical metadata makes the study more biologically interesting than prior small SCD microbiome papers. The main weakness is that the causal and mechanistic interpretation is too strong. The data support an association between SCD status and microbiome/virome features, but they do not yet establish a clear "axis of pathophysiology." The provirus findings are intriguing, but require stronger statistical control, better validation, and more cautious interpretation.

      Strengths:

      The major strengths of the study include the clinically relevant disease setting, the use of whole-community sequencing, the integration of microbial, immune-cell, cytokine, and clinical measurements, and the novel attention to bacterial virus-related features. A particularly interesting aspect of the work is the analysis of virus-like elements integrated into bacterial genomes. The authors report that these elements are enriched in the gut microbial communities of patients with sickle cell disease and are associated with several inflammatory signals in blood. This observation is potentially important because it suggests that the microbial contribution to inflammation in sickle cell disease may involve not only bacteria but also bacterial virus-related genetic elements.

      Weaknesses:

      The evidence for this proposed immune-related mechanism is incomplete. The study is cross-sectional and largely based on associations, so it cannot determine whether these virus-like elements drive immune activation, reflect immune activation, or are linked indirectly through disease severity, treatment history, or other clinical factors. The main limitations are the single-center design, modest sample size for some immune measurements, limited ability to control for treatment and disease heterogeneity, and the need for clearer multiple-testing correction in the correlation analyses. In particular, stronger adjustment for available clinical factors such as hydroxyurea use, transfusion history, pain admissions, genotype, and other markers of disease burden would help readers judge how specific the microbial and viral findings are to sickle cell disease itself.

      Overall, the authors largely achieve their descriptive aim of identifying gut microbial differences associated with sickle cell disease. The evidence is solid for the presence of broad microbial community differences, but incomplete for the stronger conclusion that virus-like elements form a pathophysiological immune axis. The work will likely be useful to researchers studying the microbiome, inflammation, and sickle cell disease, especially as a hypothesis-generating dataset. Its impact would be strengthened by more cautious interpretation, stronger control of clinical confounders, clearer statistical correction, and future longitudinal or experimental studies to test causality.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Flamholz et al. sought to determine whether consistent and significant interactions exist between the gut microbiome and disease pathology in sickle cell disease (SCD). By sequencing and analysing metagenomes from faecal samples collected from 98 SCD patients and 46 control subjects, they identified community-level shifts in both the bacterial and proviral gut microbiome of SCD patients. They further reported correlations between the proviral microbiome and multiple blood cytokines, whereas similar associations were not observed for the bacterial microbiome. Based on these findings, the authors propose the existence of a viral-immune axis in SCD pathophysiology and targetable functional alterations in the gut microbiome.

      Strengths:

      This work includes the largest SCD cohort analysed to date, enabling analysis with relatively strong statistical power. In addition to profiling the bacterial microbiome, the study also examines the gut proviral microbiome, thereby providing a more comprehensive investigation of the topic. The newly generated metagenomic dataset will also be valuable for further meta-analysis by the wider community. Overall, the authors have largely achieved their aims.

      Weaknesses:

      However, this study represents a single-centre cross-sectional investigation, and most findings remain correlative in nature. In particular, the claim that the study identifies targetable functional alterations in the gut microbiome for disease treatment may be somewhat overstated. Although the reported functional module changes in SCD patients are intriguing, additional mechanistic and/or longitudinal evidence would be required before these features can realistically be considered targetable.

    5. Author response:

      eLife Assessment:

      This study presents a valuable metagenomic analysis of the gut microbiome in sickle cell disease (SCD) patients, revealing associations between bacteriophage, host immunity, and SCD pathophysiology. While these data are interesting and helpful for hypothesis generation, they are deemed incomplete; additional experiments would be needed to test causality and to provide mechanistic insight. Despite these limitations, this work will be of broad interest to researchers studying SCD, immunology, phage biology, and the microbiome, adding to the small but growing literature suggesting a microbial component to SCD.

      The authors would like to thank the reviewers for thorough and constructive comments on our manuscript. We will make major updates to the manuscript addressing the following points and suggestions from the three reviewers: (1) assessing HbAS/AA genotype influence on microbiome composition; (2) conducting the requested beta diversity analysis, (3) conducting the requested sensitivity analysis to assess the impact of disease severity and therapy on microbiome and virome features; (4) modifying our language to clearly state that our results do not indicate causality or mechanism of microbiome interactions with sickle cell disease pathophysiology; (5) improved discussion of the phage results and their strengths and limitations; (6) additional changes throughout for clarity and correction of errors. We will change the title to “Bacterial and viral gut microbiome alterations characterize microbiome-immune-pathophysiology axes in Sickle Cell Disease.” These additions will greatly improve our work and presentation and we are grateful to the reviewers and our editors.

      We have indicated where specific changes were made in response to the public reviews below.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Flamholz and colleagues use metagenomic sequencing to profile the microbiome of individuals with sickle cell disease (SCD), the most common genetic blood disorder in the world. To build on previous studies that found dysbiosis in SCD, this manuscript aims to examine whether changes in either bacterial species or bacteriophages correlate with inflammatory hallmarks of the disease. The authors claim that sickle cell dysbiosis does not correlate with inflammatory hallmarks of the disease, but instead, aged neutrophil numbers and bacteriophages do. Appropriate control subjects and additional analyses are needed to support that conclusion.

      Strengths:

      The primary strength of this paper is the investigation into disease-associated changes in bacteriophages. This is an entirely novel idea in the sickle cell field, and based on the current results, may be an important, under-recognized disease hallmark. It is unclear, however, if phages are "the chicken or the egg" in terms of sickle cell inflammatory profiles; do these increases in phage number simply result from other disease processes, or are they in any way contributing to disease pathophysiology?

      Weaknesses:

      A primary weakness of the manuscript is the fact that the majority of individuals included in the control group maintain sickle cell trait (HbAS genotype). Although typically asymptomatic, it is unclear if this genotype is associated with microbial changes that would not be observed in a true control group (HbAA genotype). This is a significant limitation that may limit the ability to draw conclusions from the current data set.

      Another key weakness is the lack of beta diversity assessment. Although decreased alpha diversity is observed in individuals with SCD, and specific bacterial taxa are differentially abundant following multivariate analyses, there is no overall comparison of bacterial community composition between individuals with SCD and controls. Prior to drawing conclusions about the relationship (or lack thereof) between the SCD microbiome and inflammatory markers, it is important to know if this study did indeed find disease-associated changes in microbiome composition.

      It is unclear which individuals were used for aged neutrophil (AN) and molecular data assessments. For example, were children who were still receiving penicillin prophylaxis included in these specific assessments? Given the authors' previous work demonstrating that antibiotic treatment decreases AN pathology, it seems critical to limit all AN/molecular analyses to older subjects who are not on daily penicillin treatment (if possible).

      A minor weakness is the continued use of "disease" vs. "healthy" indicators as primary microbiome metrics that are used for molecular correlations. The lack of metric specificity - and lack of discussion regarding which diseases were used to generate these indicators (how similar/different are they to sickle cell?) - could be said to make these metrics essentially meaningless.

      We thank the reviewer for their helpful comments and suggestions. We want to first note that patients on prophylactic penicillin within six months of sample collection were excluded from the study due to the known impact of antibiotics on gut microbiomes, this will be clarified in the main text. We will include an analysis evaluating the influence of control genoype (HbAA/HbAS) on our microbiome and virome results. To evaluate whether control genotype influenced major microbiome and virome features, analyses were restricted to control participants only. Controls were stratified by genotype as HbAA or HbAS. Four significant microbiome and virome features were tested: F:B ratio, Shannon diversity, provirus fraction, and virus count. HbAA and HbAS controls were compared using two-sided Mann-Whitney U tests. Benjamini-Hochberg FDR correction was applied across the four tested features. HbAS and HbAA controls did not differ significantly for F:B ratio, Shannon diversity, provirus fraction, or virus count. The inclusion of HbAA/AS will strengthen our results with respect to the observation that sickle cell disease patient microbiomes remain significantly different from sickle trait (HbAS) controls. These results will be reported in a new Supplemental Table.

      We will include a beta diversity analysis using MetaPhlAn species profiles. Beta diversity analyses were performed in Python using pandas and NumPy for data processing, scikit-bio for distance calculations and PERMANOVA, scikit-learn for ordination-related computations, statsmodels for multiple-testing correction where applicable, and matplotlib for visualization.

      For the primary disease/control comparison, samples were grouped as control or SCD. For the genotype control sensitivity analysis, samples were restricted to HbAA and HbAS individuals as described above. Species detected in at least 10% of included samples were retained for beta diversity analysis. To account for the compositional structure of metagenomic relative abundance data, species profiles were transformed using a centered log-ratio transformation after addition of a small pseudocount to accommodate zero values. Aitchison distances were calculated from the CLR-transformed species profiles. Statistical significance of group separation was assessed by PERMANOVA using 999 permutations. For the control versus SCD comparison, PERMANOVA was performed between the two disease-status groups. For the HbAA versus HbAS control comparison, PERMANOVA was performed among controls only.

      In the SCD cohort, beta diversity differed significantly between controls and SCD participants by Aitchison distance after CLR transformation (R<sup>2</sup> = 0.030, p = 0.001). In contrast, HbAA and HbAS controls did not differ significantly in beta diversity (R<sup>2</sup> = 0.024, p = 0.282), supporting the conclusion that the observed SCD/control separation was not driven by control genotype composition. These methods and results will be reported in the revised manuscript.

      The manuscript describing the microbiome health and disease indicators was submitted to eLife jointly with this manuscript as a package; eLife declined to review the indicator manuscript. Briefly, this study conducted a cross-disease meta-analysis of 38 studies comprising 8,204 samples and identified 100 bacterial taxa or “indicators” that are weakly but consistently associated with health or disease across diverse conditions, including, but not limited to, inflammatory bowel disease, colorectal cancer, type 2 diabetes. The indicator taxa were validated in an independent cohort of Graves’ disease patients. We currently cite an older version of this work posted as a preprint. The manuscript is currently under review at another journal and we will update this manuscript with the updated citation when it is available.

      We will address the other recommendations from this reviewer in the revision.

      Reviewer #2 (Public review):

      Summary:

      The study analyzes stool metagenomes from 98 SCD patients and 46 controls, with SCD and control groups matched on age, race, sex, and ethnicity. The authors report lower Shannon diversity, lower Firmicutes/Bacteroidetes ratio, loss of health-associated taxa, increased disease-associated indicators, altered butyrate/fatty-acid metabolism pathways, and enrichment of provirus/prophage fractions in SCD. They further correlate aged-like neutrophils and prophage fractions with inflammatory cytokines. The strength is that this is not just another 16S comparison. The use of whole-community metagenomics, immune profiling, neutrophil assays, and clinical metadata makes the study more biologically interesting than prior small SCD microbiome papers. The main weakness is that the causal and mechanistic interpretation is too strong. The data support an association between SCD status and microbiome/virome features, but they do not yet establish a clear "axis of pathophysiology." The provirus findings are intriguing, but require stronger statistical control, better validation, and more cautious interpretation.

      Strengths:

      The major strengths of the study include the clinically relevant disease setting, the use of whole-community sequencing, the integration of microbial, immune-cell, cytokine, and clinical measurements, and the novel attention to bacterial virus-related features. A particularly interesting aspect of the work is the analysis of virus-like elements integrated into bacterial genomes. The authors report that these elements are enriched in the gut microbial communities of patients with sickle cell disease and are associated with several inflammatory signals in blood. This observation is potentially important because it suggests that the microbial contribution to inflammation in sickle cell disease may involve not only bacteria but also bacterial virus-related genetic elements.

      Weaknesses

      The evidence for this proposed immune-related mechanism is incomplete. The study is cross-sectional and largely based on associations, so it cannot determine whether these virus-like elements drive immune activation, reflect immune activation, or are linked indirectly through disease severity, treatment history, or other clinical factors. The main limitations are the single-center design, modest sample size for some immune measurements, limited ability to control for treatment and disease heterogeneity, and the need for clearer multiple-testing correction in the correlation analyses. In particular, stronger adjustment for available clinical factors such as hydroxyurea use, transfusion history, pain admissions, genotype, and other markers of disease burden would help readers judge how specific the microbial and viral findings are to sickle cell disease itself.

      Overall, the authors largely achieve their descriptive aim of identifying gut microbial differences associated with sickle cell disease. The evidence is solid for the presence of broad microbial community differences, but incomplete for the stronger conclusion that virus-like elements form a pathophysiological immune axis. The work will likely be useful to researchers studying the microbiome, inflammation, and sickle cell disease, especially as a hypothesis-generating dataset. Its impact would be strengthened by more cautious interpretation, stronger control of clinical confounders, clearer statistical correction, and future longitudinal or experimental studies to test causality.

      We thank the reviewer for their helpful comments and suggestions. We want to first note that patients on prophylactic penicillin within six months of sample collection were excluded from the study due to the known impact of antibiotics on gut microbiomes, this will be clarified in the main text. We will temper our interpretation of our results, making clear that we are not arguing that either prophages or bacteria are causal or mechanistically associated with SCD biology and pathology. We will strengthen our control of clinical confounders, and add clearer statistical correction in the revision. We look forward to conducting future studies to test causality and understand mechanism.

      We will address the other recommendations from this reviewer in the revision.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Flamholz et al. sought to determine whether consistent and significant interactions exist between the gut microbiome and disease pathology in sickle cell disease (SCD). By sequencing and analysing metagenomes from faecal samples collected from 98 SCD patients and 46 control subjects, they identified community-level shifts in both the bacterial and proviral gut microbiome of SCD patients. They further reported correlations between the proviral microbiome and multiple blood cytokines, whereas similar associations were not observed for the bacterial microbiome. Based on these findings, the authors propose the existence of a viral-immune axis in SCD pathophysiology and targetable functional alterations in the gut microbiome.

      Strengths:

      This work includes the largest SCD cohort analysed to date, enabling analysis with relatively strong statistical power. In addition to profiling the bacterial microbiome, the study also examines the gut proviral microbiome, thereby providing a more comprehensive investigation of the topic. The newly generated metagenomic dataset will also be valuable for further meta-analysis by the wider community. Overall, the authors have largely achieved their aims.

      Weaknesses:

      However, this study represents a single-centre cross-sectional investigation, and most findings remain correlative in nature. In particular, the claim that the study identifies targetable functional alterations in the gut microbiome for disease treatment may be somewhat overstated. Although the reported functional module changes in SCD patients are intriguing, additional mechanistic and/or longitudinal evidence would be required before these features can realistically be considered targetable.

      We thank the reviewer for their helpful comments and suggestions. We will note in the text that additional mechanistic and longitudinal studies are required before we can target the microbiome and virome in SCD and clarified that this is a single center, cross-sectional. We will make further modifications to the manuscript to clarify cohort features (specifically, age and race were matched, other baseline characteristics were balanced), to properly describe the Shannon diversity metric, and to fix several errors that this reviewer caught.

    1. eLife Assessment

      This is an important study that applies a new chromatin profiling technique to the study of cellular responses to low oxygen. The authors provide convincing evidence for distinct kinetic phases of the response and identify many new putative regulators of the response. This work will be of broad interest to those studying low oxygen responses and transcriptional regulation.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have satisfactorily addressed the comments raised in the previous round of review with textual revisions.]

      Summary:

      The manuscript by Singh et al. presents an application of MOA-seq to better define transcriptional control underlying the hypoxia response in human endothelial cells. This group's previously described MOA-seq technique allows for precise, identity-agnostic mapping of occupied sites of DNA-binding proteins across the epigenome and over time. Here, they applied MOA-seq to HUVECs under normal oxygen conditions or variable lengths of hypoxia treatment, comparing changes in occupancy over time and associating these changes with corresponding transcriptome alterations. This approach revealed thousands of dynamically occupied sites comprising 10 major kinetic clusters that appear to define distinct subsets and phases of the hypoxia response. Analysis of DNA motifs in these dynamically occupied regions captured the known major roles of HIF1A in the hypoxia response and also implicated new HIF1A-associated regulators. Importantly, they also identified many potential HIF1A-independent candidate TFs that act at HREs, which has been an outstanding question in the field. Additionally, this study identified ~7K additional sites not previously defined as regulatory elements by ENCODE.

      Strengths:

      Overall, this study is well executed and described, providing new biological insights as well as a rich data resource for the field. As MOA-seq was previously developed for use in plants, this work demonstrates the application of this method in mammalian cells and highlights its utility in identifying new potential regulatory sites not captured by DNase-seq or ATAC-seq. The conclusions made by the authors are well supported by the results, with the caveat that extensive use of DNA motif identification and ontology analyses invariably leads to some uncertainty regarding factor identity and gene network properties.

    3. Reviewer #2 (Public review):

      Summary:

      Singh et al. apply MOA-seq to map transcription factor occupancy genome-wide in HUVECs across a hypoxia time course. The study provides a well-validated, high-resolution view of cistrome dynamics and identifies both HIF1A-associated and independent regulatory programs.

      Major comments from the first round of review:

      Methodological validation is strong. MOA-seq's ability to map protein-bound DNA at near-nucleotide resolution without factor-specific antibodies is a genuine advance, and the cross-validation against independent ChIP-seq and ENCODE datasets is convincing. As noted, future work with additional biological replicates could further strengthen confidence in the smaller kinetic clusters.

      Imaging-based validation would strengthen the key biological claims. The kinetic clustering and pathway enrichments are computationally inferred. Orthogonal approaches, for example, live-cell fluorescence imaging of HIF1A nuclear translocation to confirm the proposed temporal binding waves, would provide independent experimental support.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Singh et al. presents an application of MOA-seq to better define transcriptional control underlying the hypoxia response in human endothelial cells. This group's previously described MOA-seq technique allows for precise, identity-agnostic mapping of occupied sites of DNA-binding proteins across the epigenome and over time. Here, they applied MOA-seq to HUVECs under normal oxygen conditions or variable lengths of hypoxia treatment, comparing changes in occupancy over time and associating these changes with corresponding transcriptome alterations. This approach revealed thousands of dynamically occupied sites comprising 10 major kinetic clusters that appear to define distinct subsets and phases of the hypoxia response. Analysis of DNA motifs in these dynamically occupied regions captured the known major roles of HIF1A in the hypoxia response and also implicated new HIF1A-associated regulators. Importantly, they also identified many potential HIF1A-independent candidate TFs that act at HREs, which has been an outstanding question in the field. Additionally, this study identified ~7K additional sites not previously defined as regulatory elements by ENCODE.

      Strengths:

      Overall, this study is well executed and described, providing new biological insights as well as a rich data resource for the field. As MOA-seq was previously developed for use in plants, this work demonstrates the application of this method in mammalian cells and highlights its utility in identifying new potential regulatory sites not captured by DNase-seq or ATAC-seq. The conclusions made by the authors are well supported by the results, with the caveat that extensive use of DNA motif identification and ontology analyses invariably leads to some uncertainty regarding factor identity and gene network properties.

      Weaknesses:

      There are several areas where the clarity of presentation could be improved:

      (1) Given the importance of the methodology, the methods section needs more detail on how the extent of MNase digestion is chosen to achieve optimal results with MOA-seq. This is described to some extent in the description of control library preparation, but not for the experimental samples.

      We thank the reviewer for noting this unintended omission. We have not updated the Methods section to specify as follows:

      "Digestion patterns were assessed via gel electrophoresis, and the light digest levels ideal for MOA-seq (as per Savadel et al., 2021) were selected as the lightest digest levels that give a pattern of a nucleosomal ladder spanning the entire DNA fragment size range from undigested to mononucleosome bands, as indicated in Figure 1 with the asterisk-marked gel lanes."

      (2) The abstract describes this approach as "native cistrome profiling" but this is misleading since formaldehyde fixation is used.

      We believe the formaldehyde fixation captures native chromatin structure, but indeed we are digesting fixed chromatin and have updated the wording to read as “in situ cistrome profiling.”

      (3) Species- and field-specific jargon and abbreviations need to be clarified on first usage. For example, on page 9: "Downsampling analysis was carried out for two sets of published reference peaks; the CTCF cCRE peak midpoints and for the ERG motif under the ERG ReMap ChIP-seq peaks." The different categories of cCREs were not clearly defined, nor will it be clear what the term ReMap refers to for those outside the field. The sentence after this refers to IDR, which also should be defined.

      We thank the reviewer for highlighting the need for clearer definitions of field-specific terminology and abbreviations. In response, we have revised the manuscript to explicitly define all relevant terms at first mention. Specifically, we now describe the ENCODE candidate cis-regulatory element (cCRE) catalogue and define the individual cCRE categories, including promoter-like (PLS), proximal enhancer-like (pELS), distal enhancer-like (dELS), DNase I–H3K4me3 (K4m3), and CTCF-only regions. We also clarify that ReMap is a curated database of human transcriptional regulator binding peaks derived from ChIP-seq, ChIP-exo, and DAP-seq experiments. Additionally, we now define IDR as the Irreproducible Discovery Rate framework upon first use.

      (4) Figure 4C: Are these motifs examined under MOA sites specifically or anywhere in the genes in question?

      Leading up to and including Figure 4C, we have not yet examined any motifs. Instead, Figure 4C compares gene sets, one defined by our diff-MOA, and those from GO libraries, in this case the "target genes" which are defined by TF-specific studies, primarily ChIP-seq but also related immuno-based mapping techniques. Consequently, the analysis shown in Fig. 4C is not a motif enrichment analysis. Instead, we used the ENRICHR gene set enrichment analysis tool with ENCODE and ChEA consensus transcription factor target gene sets. Thus, the analysis was performed at the gene-set level, and transcription factor motifs were not examined within diff-MOA peaks or elsewhere in the associated genes for Fig. 4C. We note that motif enrichment within diff-MOA peaks was subsequently examined separately in Fig. 6. In Fig. 7, we further examined differentially expressed genes associated with diff-MOA peaks containing enriched transcription factor motifs and used clustering analyses to investigate their regulatory relationships. We have clarified these distinctions in the revised manuscript.

      If the question is about the location of MOA footprints relative to gene structure, we did not examine any MOA sites at any specific location, just overlapping the gene +/- 200 bp, as indicated in Fig. 4B.

      (5) Figure 5B shows that up-DEGs with diff-MOA footprints tend to show more losses of footprints. Do the authors interpret this as a loss of repressor binding?

      Not exclusively, but yes, that is one plausible explanation. That is, the activation (defined by increased RNA levels) via de-repression could be happening. But we also expect these dynamic footprints to be but one component. In other words, we interpret the relationship as consistent with that possibility, but not only that possibility. A logical explanation is that loss of footprint occupancy associated with upregulated genes could be based on displacement of repressive DNA-binding factors, thereby contributing to transcriptional activation. Thus, while loss of repressor binding is a plausible explanation for a subset of these events, additional factor-specific experiments would be required to know for sure in each case. We have added text to the Discussion acknowledging this possibility.

      Reviewer #2 (Public review):

      Summary:

      Singh et al. apply MOA-seq to map transcription factor occupancy genome-wide in HUVECs across a hypoxia time course. The study provides a well-validated, high-resolution view of cistrome dynamics and identifies both HIF1A-associated and independent regulatory programs.

      Major Comments:

      Methodological validation is strong. MOA-seq's ability to map protein-bound DNA at near-nucleotide resolution without factor-specific antibodies is a genuine advance, and the cross-validation against independent ChIP-seq and ENCODE datasets is convincing. As noted, future work with additional biological replicates could further strengthen confidence in the smaller kinetic clusters.

      Regarding additional biological replicates, we have acknowledged this point in the discussion. Importantly, we did subject the replicates to IDR analysis, which we explain in the methods as "In accordance with ENCODE ChIP-seq guidelines (Landt et al., 2012), we further evaluated data quality by assessing pooled pseudo-replicate consistency and self-consistency for each individual replicate (Supplementary Table S2)." This IDR analysis demonstrated consistent peaks between our bioreplicates, meeting ENCODE guidelines. In addition, downsampling analysis demonstrated that our sequencing depth of coverage (Supp Fig 1) was over 10-fold greater than required. We do appreciate that it will be useful to have more biological replicates from other cell types, tissues, or organisms, and hope this study prompts just such future research.

      Imaging-based validation would strengthen the key biological claims. The kinetic clustering and pathway enrichments are computationally inferred. Orthogonal approaches, for example, live-cell fluorescence imaging of HIF1A nuclear translocation to confirm the proposed temporal binding waves, would provide independent experimental support.

      Live-cell imaging could indeed be interesting, but it is beyond our current capacity to add to this study and consider this an exciting future direction, but presence in the nucleus could include both bound and unbound HIF1, so the results may not easily track the DNA-bound HIF1 only.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      In Figure 3B, the x-axis is not labeled.

      Thank you for pointing this out. We have revised Figure 3B by adding the previously missing x-axis label.

      Reviewer #2 (Recommendations for the authors):

      In the abstract, it would be good to define what MOA-seq is and what the cistrome is.

      Thank you for this suggestion. We have revised the abstract to define both MOA-seq (MNase-defined cistrome-Occupancy Analysis sequencing) and the cistrome upon first mention to improve accessibility for readers who may be unfamiliar with these terms.

    1. eLife Assessment

      The manuscript concerns a fundamental and controversial question in Trypanosoma brucei biology and the parasite life cycle, whether or not dividing slender bloodstream forms must transition to growth-arrested stumpy forms before differentiating to the procyclic form in the Tsetse midgut. The authors provide further evidence that slender bloodstream forms can infect Tsetse flies, and that although their differentiation is considerably delayed, they do not become classical stumpy forms during the process. The study is solid in design and execution, and addresses several criticisms made of the authors' earlier work, although discrepancies with results from other laboratories remain.

    2. Reviewer #2 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers.]

      Summary:

      This paper is an exciting follow-up to two recent publications in eLife: one from the same lab, reporting that slender forms can successfully infect tsetse flies (Schuster, S et al., 2021), and another independent study claiming the opposite (Ngoune, TMJ et al., 2025). Here, the authors address four criticisms raised against their original work: the influence of N-acetyl-glucosamine (NAG), the use of teneral and male flies, and whether slender forms bypass the stumpy stage before becoming procyclic forms.

      Strengths:

      We applaud the authors' efforts in undertaking these experiments and contributing to a better understanding of the T. brucei life cycle. The paper is well-written and the figures are clear.

      Comments on revisions:

      We thank the authors for the revised manuscript and for considering our comments.

    3. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      Summary:

      This paper is an exciting follow-up to two recent publications in eLife: one from the same lab, reporting that slender forms can successfully infect tsetse flies (Schuster, S et al., 2021), and another independent study claiming the opposite (Ngoune, TMJ et al., 2025). Here, the authors address four criticisms raised against their original work: the influence of N-acetyl-glucosamine (NAG), the use of teneral and male flies, and whether slender forms bypass the stumpy stage before becoming procyclic forms.

      Strengths:

      We applaud the authors' efforts in undertaking these experiments and contributing to a better understanding of the T. brucei life cycle. The paper is well-written and the figures are clear.

      Comments on revisions:

      We thank the authors for the revised manuscript and for considering our comments.

      We outline below the 3 points that, in our opinion, remain to be clarified.

      (1) Effect of NAG on slender-form infections in tsetse flies

      The conclusion that "NAG has a negligible effect on slender infections in tsetse flies" based on Figure 1, cannot be fully supported in the absence of a positive control. A relevant positive control is well established in the literature, namely that NAG promotes Tsetse infection by stumpy forms. Without such a control, it is not possible to exclude technical issues (for example, an ineffective NAG treatment), which would yield results similar to those presented in Figure 1.

      We agree that an internal stumpy-form positive control would provide an additional technical reference. However, the enhancing effect of NAG on stumpy-form midgut infections is well established and was also demonstrated under the experimental framework of our original study (Schuster et al. 2021, Figure 2A).

      The purpose of the present Research Advance was therefore not to re-establish the known effect of NAG on stumpy infections, but to test whether slender-form infections require NAG supplementation. Under the conditions tested here, slender bloodstream forms established midgut, proventriculus and salivary-gland infections also in the absence of NAG. We have revised the text accordingly to avoid implying a general absence of NAG effects and to make clear that our conclusion is restricted to slender-form infections under the conditions tested (line 128).

      (2) Infection of non-teneral flies

      Because the experiments shown in Figure 1 (teneral flies) and Figure 2 (non-teneral flies) were not conducted in parallel or under identical conditions, it is important that the figure legends clearly state the parasite numbers used in each case. Specifically, infections of teneral flies were performed with 200 parasites/mL (approximately 4 parasites per bloodmeal), whereas non-teneral infections used 1 × 10<sup>6</sup> parasites/mL (approximately 20,000 parasites per bloodmeal?). At present, this information is scattered across the Methods and Supplementary Tables 1 and 2, making it difficult for readers to immediately appreciate that the parasite load differs by roughly 5,000-fold between these conditions.

      As previously shown by the authors (Schuster et al., 2021) and in the Rotureau laboratory (Tsagmo Ngoune et al.), and as generally expected, the initial parasite dose strongly influences infection outcomes in teneral flies. In this context, it would be informative to know whether the authors have attempted infections of non-teneral flies using lower parasite numbers (noting that Tsagmo Ngoune et al. used a maximum of 10,000 parasites) and what the infection rate was.

      Relatedly, the statement in line 370 appears to be an overgeneralization, as fly age was not directly tested under matched experimental conditions:

      Line 370 - "Here, we unambiguously show that, in the absence of immunosuppressive treatment, slender forms can establish infections in tsetse flies, irrespective of the fly's age or sex."

      We thank the reviewer for highlighting the inconsistent presentation of parasite doses between Figure 1 and 2. We agree this is confusing and have revised the figure legends to clearly state both the parasite concentration (cells/mL) and estimated fly uptake per bloodmeal for each experiment (Lines 143 and 206).

      Regarding experiments with non-teneral flies using lower parasite numbers: We have not tested intermediate doses (e.g., 10,000 parasites/bloodmeal as used by Ngoune et al.) in non-teneral flies. Given that teneral flies already show relatively low infection rates even under optimal conditions, we chose the higher parasite dose (20,000 parasites/bloodmeal) for non-teneral flies to ensure sufficient statistical power for meaningful analysis of infection outcomes across different fly compartments.

      We acknowledge the reviewer's concern regarding the statement in line 370 and have revised this sentence (line 375) to more accurately reflect our experimental conditions, avoiding overgeneralization beyond the specific parameters tested.

      This reads now: “Here, we demonstrate that slender forms can establish infections without immunosuppressive treatment under the conditions tested. This infectivity was observed in both teneral and non-teneral, as well as in both male and female flies, indicating that slender forms retain transmission potential across different fly demographics. However, direct age comparisons under identical parasite doses remain to be tested.”

      (3) Transcriptomic analysis

      Supplementary Figure 8 lacks statistical analysis, which limits its interpretability. Two types of comparisons would be particularly helpful:

      (i) a comparison of PAD1/2 expression levels between slender and stumpy forms at 0 h; and

      (ii) for each gene, a comparison of the overall change in expression (from 0 to 72 h) between infections initiated with slender versus stumpy forms.

      In addition, the figure legend should clarify what "expression levels" refer to. TPM? Normalized counts?

      We appreciate this helpful comment and included statistical analysis for the expression of PAD1 and PAD2 (Supplementary Figure 8) between the two forms for the baseline (0 h) as well as during the differentiation to procyclic forms (0 h to 72 h) by using Welch´s t-test.

      While PAD1 did not show a statistically significant difference in this analysis, PAD2 displayed significant differences in expression dynamics over time. This supports the broader transcriptomic observation that slender- and stumpy-initiated differentiation follow distinct transcriptional trajectories before converging at the procyclic stage.

      We also clarified the figure legends showing the mean log2 counts per million (CPM) values.

      Finally, for the benefit of the field, eLife could encourage publishing a collaborative study in which the Engstler and Rotureau laboratories exchange parasite lines and culture protocols (including media with and without methylcellulose) and perform tsetse fly infections in parallel in their respective laboratories. Such an approach could help resolve the remaining discrepancies and provide a valuable reference for the community.

      We appreciate this constructive suggestion. A collaborative inter-laboratory study in which parasite lines, culture conditions and infection protocols are exchanged between the Engstler and Rotureau laboratories would be a valuable way to address the remaining discrepancies in the field. In particular, parallel infections using matched parasite lines and culture conditions, including media with and without methylcellulose, could provide a useful reference dataset for the community.

      At the same time, such a study would require substantial coordination, reciprocal strain exchange, protocol harmonization and new infection series in two laboratories. It therefore goes beyond the scope of the present Research Advance, which was designed to address the specific methodological concerns raised in response to our original publication. We have restricted our conclusions accordingly and view the proposed collaborative benchmark study as an important direction for future work.

    1. eLife Assessment

      This important study reports the development of the first tankyrase degrader and demonstrates its enhanced ability to inhibit β-catenin signaling compared to conventional tankyrase inhibitors. The evidence supporting the conclusions is comprehensive and convincing, based on rigorous biochemical and cellular analyses. The findings will be of broad interest to researchers studying Wnt signaling, protein degradation, and cancer biology.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript reports the discovery and characterization of the first bifunctional degrader of tankyrase. Notably, the tankyrase degrader exhibits stronger β-catenin inhibition and tumor growth suppression compared to conventional tankyrase inhibitors. Mechanistically, while tankyrase inhibitors stabilize tankyrase and promote Axin puncta formation-thereby impairing β-catenin degradation-the degrader avoids this effect, resulting in deeper suppression of β-catenin signaling. These findings suggest that targeted degradation of tankyrase offers a novel therapeutic strategy for β-catenin-driven cancers. Overall, this is a compelling study with significant translational potential.

      Strengths:

      (1) The manuscript presents a rigorous and well-executed study on a timely and impactful topic.

      (2) The biochemical and cellular characterization of the tankyrase degrader is thorough, and the comparative analysis with tankyrase inhibitors is insightful.

      (3) The finding that tankyrase stabilization by inhibitors may interfere with Axin function is novel and significant. It aligns with earlier observations (e.g., Huang 2009) that transient tankyrase overexpression can stabilize β-catenin independently of PAR domain activity.

      (4) The use of TNKS1/2 knockout cells expressing catalytically inactive tankyrase to demonstrate β-catenin inhibitory activity of the tankyrase degrader is elegant.

      (5) The finding that the tankyrase degrader has superior anti-proliferative effects in colorectal cancer models has important therapeutic implications.

      Comments on revised version:

      I had a favorable opinion of the manuscript in the first round of review. I don't have additional comments on the revised manuscript. The manuscript looks fine to me.

    3. Reviewer #2 (Public review):

      Summary:

      The ADP-ribosyltransferase tankyrase controls many biological processes, many of which are relevant to human disease. This includes Wnt/beta-catenin signalling, which is dysregulated in many cancers, most notably colorectal cancer. Tankyrase is a positive regulator of Wnt/beta-catenin signalling in that it counters the activity of the beta-catenin destruction complex (DC). Catalytic inhibition of tankyrase not only blocks PAR-dependent ubiquitylation and degradation of AXIN1/2, the central scaffolding protein in the DC, but also tankyrase itself. As a result, blocking tankyrase gives rise to tankyrase accumulation, which may accentuate its non-catalytic functions, which have been proposed to drive Wnt/beta-catenin signalling. Most tankyrase catalytic inhibitors have shown limited efficacy and substantial toxicity in vivo. By developing tankyrase-directed PROTACs, the authors aim to block both catalytic and non-catalytic functions of tankyrase, aspiring to achieve a more complete inhibition of Wnt/beta-catenin signalling. The successfully developed PROTAC, based on the existing catalytic inhibitor IWR1, IWR1-POMA, induces the degradation of both TNKS and TNKS2, blocks beta-catenin-dependent transcription without stabilising the DC in puncta/degradasomes, and inhibits cancer cell growth in vitro. Mechanistically, this points to a scaffolding role of tankyrase in the DC, at least under conditions of tankyrase catalytic inhibition, in line with previous proposals.

      Strengths:

      The study clearly illustrates the incentive for developing a tankyrase degrader, namely, to abolish both catalytic and non-catalytic functions of tankyrase. By and large, the study achieves these ambitions, and the findings support the main conclusions, although the statement that a more complete inhibition of the pathway is achieved requires corroboration. The proteomics studies are powerful. IWR1-POMA constitutes a very useful tool to re-evaluate targeting of tankyrase in oncogenic Wnt/beta-catenin signalling. The paired compounds will benefit investigations of tankyrase scaffolding functions across many different biological systems controlled by tankyrase. The findings are exciting.

      Comments on revised version:

      I thank the authors for responding to the queries raised in the original review, most of which have now been addressed. This further strengthens this well-conducted study and well-presented manuscript. I congratulate the authors for this interesting and insightful work.

      A few minor points remain:

      I appreciate the authors acknowledge that testing the physical properties of the degradasome puncta is necessary to explore whether they indeed represent condensates. The term "condensates" implies liquid-liquid phase separation (rightly or wrongly). However, this question has not yet been resolved in the case of degradasomes. I therefore suggest the term "condensates" to be avoided. A simple morphological description as "puncta" may suffice.

      I thank the authors for including the additional data comparing tankyrase binding by IWR and IWR-POMA. I agree that using the BRET signal of IWR-POMA is informative. Adding the IC50 values directly to the figure panels (S3E, S3G) would help the reader to quickly assess binding. The comparison between these two panels is insightful.

      Regarding the use of the terms TNKS, TNKS1 and TNKS2, if the authors would like to use the name "TNKS" to refer to both paralogues collectively, can this please be specified early in the manuscript to limit confusion with the official gene name "TNKS", which of course only refers to one paralogue?

    4. Reviewer #3 (Public review):

      In this manuscript, Wang et al employ a chemical biology approach to investigate the differences between the enzymatic and scaffolding roles of tankyrase during Wnt β-catenin signalling. It was previously established that, in addition to its enzymatic activity, tankyrase 1/2 also plays a scaffolding function within the destruction complex, a property conferred by SAM-domain-dependent polymerization (PMID: 27494558). It is also known that TNKS1/2 is an autoregulated protein and that its enzymatic inhibition leads to accumulation of total TNKS proteins and stabilization of Axin punctae (through the scaffolding function of TNKS1/2), leading to rigidification of the DC and decreased β-catenin turnover. The authors surmised that this could, in part, explain the limited efficacy of TNKS1/2 catalytic inhibition for the treatment of colorectal cancers. To test this hypothesis, they evaluated a series of PROTAC molecules promoting the degradation of TNKS1/2 to block both the catalytic and scaffolding activities. They show that IWR1-POMA (their most active molecule) promotes more efficient suppression of beta-catenin-mediated transcription and is more active in inhibiting colorectal cancer cell and CRC patient-derived organoids growth. Mechanistically, the authors used FRAP to demonstrate that catalytic inhibitors of TNKS led to a reduced dynamic assembly of the DC (rigidification), whereas IWR1-POMA did not affect the dynamics.

      Overall, this is an interesting study describing the design and development of a PROTAC for TNKS1/2 that could have increased efficacy where catalytic inhibitors have displayed limited activity. Knowing the importance of the scaffolding role of TNKS1/2 within the destruction complex, targeting both the catalytic and scaffolding roles certainly makes sense. The manuscript contains convincing evidence of the different mechanisms of the PROTAC vs catalytic inhibitors. Some additional efforts to quantify several of the experiments and to indicate the reproducibility and statistical analysis would strengthen the manuscript. Ultimately, it would have been great to evaluate the in vivo efficacy of IWR1-POMA in an in vivo CRC assay (APCmin mice or using PDX models); however, I realize that this is likely beyond the scope of this manuscript.

    5. Reviewer #4 (Public review):

      From the Reviewing Editor:

      This important study reports the development of the first PROTACs targeting the ADP-ribosyltransferases tankyrase 1 and 2, with the goal of inhibiting Wnt/β-catenin signaling more completely than is possible with catalytic tankyrase inhibitors. The work addresses a significant limitation of existing tankyrase inhibitors: although catalytic inhibition stabilizes AXIN1/2 and suppresses Wnt signaling, it also stabilizes tankyrase itself, potentially enhancing non-catalytic scaffolding functions and promoting accumulation of degradasome-like puncta.

      The evidence is convincing. The authors use appropriate and well-validated approaches, including chemical biology, cellular assays, and proteomic profiling, to show that PROTAC-mediated degradation of tankyrase avoids tankyrase accumulation while still stabilizing AXIN and inhibiting Wnt/β-catenin signaling. The data support the conclusion that degradation of tankyrase can separate pathway inhibition from the confounding effects of stabilized tankyrase protein and may therefore offer advantages over conventional catalytic inhibitors.

      A strength of the study is the clear mechanistic comparison between tankyrase degradation and catalytic inhibition. The manuscript provides convincing evidence that the PROTAC and catalytic inhibitors act through distinct mechanisms, with the PROTAC targeting both catalytic and scaffolding roles of tankyrase. The study is well conducted and clearly presented, and the authors have addressed most concerns raised during review.

      A remaining limitation is that the therapeutic potential of the compound is not tested in vivo, for example in APC-mutant colorectal cancer models, APCmin mice, or patient-derived xenografts. Such experiments would strengthen claims about practical efficacy, although they are not essential for the main mechanistic conclusions of the manuscript.

      Overall, this is an important and insightful contribution. It advances the tankyrase and Wnt signaling fields by providing a new chemical strategy to suppress tankyrase function more completely than catalytic inhibition alone, and it offers a useful framework for future therapeutic exploration of tankyrase degradation.

    6. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript reports the discovery and characterization of the first bifunctional degrader of tankyrase. Notably, the tankyrase degrader exhibits stronger β-catenin inhibition and tumor growth suppression compared to conventional tankyrase inhibitors. Mechanistically, while tankyrase inhibitors stabilize tankyrase and promote Axin puncta formation - thereby impairing β-catenin degradation - the degrader avoids this effect, resulting in deeper suppression of β-catenin signaling. These findings suggest that targeted degradation of tankyrase offers a novel therapeutic strategy for β-catenin-driven cancers. Overall, this is a compelling study with significant translational potential.

      Strengths:

      (1) The manuscript presents a rigorous and well-executed study on a timely and impactful topic.

      (2) The biochemical and cellular characterization of the tankyrase degrader is thorough, and the comparative analysis with tankyrase inhibitors is insightful.

      (3) The finding that tankyrase stabilization by inhibitors may interfere with Axin function is novel and significant. It aligns with earlier observations (e.g., Huang 2009) that transient tankyrase overexpression can stabilize β-catenin independently of PAR domain activity.

      (4) The use of TNKS1/2 knockout cells expressing catalytically inactive tankyrase to demonstrate β-catenin inhibitory activity of the tankyrase degrader is elegant.

      (5) The finding that the tankyrase degrader has superior anti-proliferative effects in colorectal cancer models has important therapeutic implications.

      Weaknesses:

      (1) A key caveat is that the identified tankyrase degrader also targets GSPT1 for degradation. This raises the possibility that GSPT1 degradation may contribute to the observed β-catenin and tumor growth inhibition.

      (2) The authors address this concern reasonably by showing that DLD1 cells resistant to GSPT1 degradation remain sensitive to the tankyrase degraded.

      (3) To further strengthen this point, the authors might consider generating TNKS1/2 double knockout cells (e.g., in DLD1 or SW480 backgrounds) and demonstrating that the degrader loses its growth-inhibitory effect in these models. However, given the technical challenges of creating double knockouts in cancer cell lines, such experiments could be considered optional.

      We thank the Reviewer for the favorable feedback. The major concern is the collateral degradation of GSPT1. As the Reviewer noted, IWR1-POMA was able to suppress colony formation in DLD-1 cells resistant to a GSPT1/2 degrader (DLD-1R, Figure 6B and S9F), suggesting that TNKS but not GSPT degradation is responsible for growth inhibition.

      We also appreciate that the Reviewer brought it to our attention an important early observation of the TNKS scaffolding effects. Cong reported in 2009 that overexpression of TNKS induced AXIN puncta formation in a SAM but not PARP domain-dependent manner (PMID: 19759537, Ref. 12). We have added this reference to the introduction of TNKS scaffolding in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      The ADP-ribosyltransferase tankyrase controls many biological processes, many of which are relevant to human disease. This includes Wnt/beta-catenin signalling, which is dysregulated in many cancers, most notably colorectal cancer. Tankyrase is a positive regulator of Wnt/beta-catenin signalling in that it counters the activity of the beta-catenin destruction complex (DC). Catalytic inhibition of tankyrase not only blocks PAR-dependent ubiquitylation and degradation of AXIN1/2, the central scaffolding protein in the DC, but also tankyrase itself. As a result, blocking tankyrase gives rise to tankyrase accumulation, which may accentuate its non-catalytic functions, which have been proposed to drive Wnt/beta-catenin signalling. Most tankyrase catalytic inhibitors have shown limited efficacy and substantial toxicity in vivo. By developing tankyrase-directed PROTACs, the authors aim to block both catalytic and non-catalytic functions of tankyrase, aspiring to achieve a more complete inhibition of Wnt/beta-catenin signalling. The successfully developed PROTAC, based on the existing catalytic inhibitor IWR1, IWR1-POMA, induces the degradation of both TNKS and TNKS2, blocks beta-catenin-dependent transcription without stabilising the DC in puncta/degradasomes, and inhibits cancer cell growth in vitro. Mechanistically, this points to a scaffolding role of tankyrase in the DC, at least under conditions of tankyrase catalytic inhibition, in line with previous proposals.

      Strengths:

      The study clearly illustrates the incentive for developing a tankyrase degrader, namely, to abolish both catalytic and non-catalytic functions of tankyrase. By and large, the study achieves these ambitions, and the findings support the main conclusions, although the statement that a more complete inhibition of the pathway is achieved requires corroboration. The proteomics studies are powerful. IWR1-POMA constitutes a very useful tool to re-evaluate targeting of tankyrase in oncogenic Wnt/beta-catenin signalling. The paired compounds will benefit investigations of tankyrase scaffolding functions across many different biological systems controlled by tankyrase. The findings are exciting.

      Weaknesses:

      Although the results are promising and mostly compelling, the claim that the PROTACs provide "a deeper suppression of the WNT/β-catenin pathway activity" requires further corroboration, particularly at endogenous tankyrase levels.

      We thank the Reviewer for the encouraging and insightful comments. The major critique concerns whether TNKS degraders can suppress WNT/β-catenin signaling more effectively than TNKS inhibitors at endogenous TNKS levels. IWR1-POMA reduced the level of cytosolic β-catenin more effectively than IWR1 in Wnt3A-stimulated HEK293 cells without protein overexpression (Figure 1D). IWR1POMA also suppressed STF activity more effectively than IWR1 in DLD-1 cells (Figure S8C) and reduced the expression levels of several WNT/β-catenin targets more effectively than IWR1 (Figure 1G and S8D). These results support that TNKS degraders can suppress WNT/β-catenin signaling more effectively than TNKS inhibitors at endogenous TNKS levels.

      There are also some other points that, if considered, would further improve the manuscript, as detailed below.

      (1) Abstract and line 62: Many catalytic tankyrase inhibitors tend to display toxicity, which is likely on-target (e.g., 10.1177/0192623315621192; 10.1158/0008-5472). This constitutes the main limiting factor for these compounds. An incomplete inhibition of Wnt/beta-catenin signalling may contribute to the challenges, but this does not appear to be the dominant problem. A more prominent introduction to this important challenge is probably expected by the field.

      A previous study showed that G007-LK, a selective TNKS inhibitor, exhibited weak efficacy and dose-limiting toxicity at 5‒30 mg/kg BID or 10‒60 mg/kg QD in various mouse xenograft models (PMID: 23539443, Ref. 28). Similarly, G-631, another TNKS inhibitor, also showed dose-limiting toxicity without significant efficacy at 25‒100 mg/kg QD in mice (PMID: 26692561, Ref. 60). However, other studies showed that G007-LK was well-tolerated at 200 mg/kg QD over 3 weeks in mice (PMID: 29316982, Ref. 61), and treating mice with G007-LK at 10 mg/kg QD over 6 months also improved glucose tolerance without notable toxicity (PMID: 26631215, Ref. 62). Importantly, basroparib, a selective TNKS inhibitor, was well tolerated in a recent clinical trial (PMID: 40964966, Ref. 64), and constitutive silencing of both TNKS1 and TNKS2 for 150 days in APC-null mice prevented tumorigenesis without damaging the intestines (PMID: 31337618, Ref. 8). We have included some discussion of the toxicity issue associated with TNKS targeting at the end of the Discussion section.

      (2) The authors do a good job in setting the scene for the need for tankyrase degraders. Their observations relating to the formation of puncta (degradasomes) being tankyrase-dependent are compatible with a previous study by Martino-Echarri et al. 2016 (10.1371/journal.pone.0150484): simultaneous silencing of TNKS and TNKS2 by RNAi abolishes degradasome formation. The paper is cited as reference 17, but only in passing, and deserves more prominence. (It includes an entire paragraph titled "Expression of tankyrases 1 and 2 is required for TNKSi-induced formation of axin puncta").

      Indeed, Henderson’s 2016 paper (PMID: 26930278, previously Ref. 17, now Ref. 18) shed important light on the role of TNKS scaffolding in the DC. However, whereas this study demonstrated that knocking down both TNKS1 and TNKS2 by siRNA prevented G007-LK to induce AXIN puncta, it concluded that “puncta formation requires both the expression and the inactivation of TNKS,” which is inconsistent with our observations that accumulation of either catalytically active or inactive TNKS can promote AXIN puncta formation. The function roles of TNKS scaffolding in the DC also remained unaddressed. We have included additional discussion of Henderson’s findings in the first paragraph the Discussion section.

      (3) Moreover, the scaffolding concept has been discussed comprehensively in other studies: 10.1111/bph.14038 and more recently 10.1042/BCJ20230230. There are also a few studies that focus on targeting the ankyrin repeat clusters of tankyrase to disengage substrates (10.1038/s41598-020-69229-y; 10.1038/s41598-019-55240-5) that illustrate the concept of blocking the scaffolding function. In that sense, the hypotheses are mature, and it is interesting to see some of them supported in this study. The authors could improve how they set their work into the context of these other efforts and proposals.

      Indeed, Guettler demonstrated in 2016 that TNKS scaffolding could promote WNT/β-catenin signaling, which forms the basis of the current work. Meanwhile, whereas there have been efforts to target the SAM or ARC domain to address TNKS scaffolding by Guettler and Lehtiö, our approach of targeting TNKS for degradation is complementary. We have included in the last paragraph of the Discussion section information on efforts to target the ARC or SAM domains as an alternative approach to suppress WNT/β-catenin signaling without promoting TNKS oligomerization (PMID: 31836723 and 32704068, Ref. 66 and 67).

      (4) In several places in the manuscript, the DC is referred to as "biomolecular condensate", at times even as a "classic example", implying that it operates through phase separation. This has not been demonstrated. In fact, super-resolution microscopy indicates that the puncta are not droplet-like (10.7554/eLife.08022), which would argue against the condensate hypothesis.

      Biomolecular condensates are membraneless cellular compartments formed by phase separation of biomolecules, regardless of their physical/material properties (PMID: 28935776 and 28225081, Ref. 22 and 23). Super-resolution microscopy studies by Stenmark (PMID: 26124443, Ref. 17) showed that AXIN, APC, TNKS, and β-catenin interacted with each other to assemble into membraneless complexes, wherein AXIN and APC formed filaments throughout the DC. Peifer has also summarized evidence that supports the condensate nature of the DC (PMID: 30782412, Ref. 9; see also PMID: 26393419). However, we acknowledge that testing the physical properties of reconstituted DC (for example, PMID: 34352208) with TNKS will provide a better understanding of the nature, for example liquid vs. gel, of these condensates.

      (5) It is beautiful to be able to use IWR1 and IWR1-POMA at identical concentrations for direct comparisons. However, this requires the two compounds to bind to tankyrase similarly well and reach the target to a comparable extent. How sure are authors that target engagement is comparable? Has this been evaluated?

      Using a BRET assay, we have confirmed that IWR1-POMA binds to TNKS1 with affinity comparable to that of IWR1. Details of this study is now included in the Results sections, and the data are presented in the Supplementary Information (Fig. S3E–G).

      (6) Figure 1F: It is not immediately apparent how IWR1-POMA shows more complete containment of Wnt/beta-catenin signalling. Most Wnt/beta-catenin targets lie close to the perfect diagonal, so I do not see how the statement "that IWR1-POMA controlled WNT/β-catenin signaling more effectively than IWR1" (in the legend of Figure 1F) is supported. Minimally, an expanded explanation would benefit the reader. Providing the colour-coding legend directly in the figure would help improve clarity. Also, the panel is very small and may benefit from a different presentation in the figure.

      We have updated Fig. 1F to include an inset of Quadrant III for improved clarity and readability. We have also moved Fig. S7C to the main text as Fig. 1G and added an expanded explanation for these figures.

      (7) Figure 2: The conclusion of a "deeper suppression" of signalling relies on overexpression of tankyrase in an otherwise tankyrase-null background. Have the authors attempted to measure reporter activity or endogenous gene expression without tankyrase overexpression, in Wnt3a-stimulated cells (in the context of a normal Wnt/beta-catenin pathway) or CRC cells at the basal level? Non-catalytic activity in a similar assay has previously been observed upon tankyrase overexpression (10.1016/j.molcel.2016.06.019). Whether or not there is a substantial scaffolding effect at endogenous tankyrase levels after tankyrase inhibition remains unconfirmed, and the PROTAC is a valuable tool to address this important question. The findings presented in Figure S7C and D go some way towards answering this question - these data could be presented more prominently, and similar assays could be performed in other cell systems.

      IWR1-POMA suppressed STF activity more effectively than IWR1 in APC-mut DLD-1 and SW480 CRC cells without TNKS overexpression (Fig. S8C). Similarly, IWR1-POMA provided a deeper suppression of STF signals in HeLa cells transfected with AXIN1 and β-catenin while expressing endogenous TNKS (Fig. 4G). These results suggest that inhibitor-induced TNKS scaffolding plays a significant role at endogenous TNKS expression levels. Following the reviewer’s suggestion, Fig. S7C is now Fig. 1G.

      (8) Line 237/238: "TNKS accumulation negatively impacts the catalytic activity of the DC (Figure 5D)" - the data do not show this. Beta-catenin levels are a surrogate readout for DC function (phosphorylation and ubiquitylation). Minimally, this requires rewording, with reference to beta-catenin levels.

      We have rephrased "TNKS accumulation negatively impacts the catalytic activity of the DC" as "TNKS accumulation negatively impacts the exchange of β-catenin in the DC."

      (9) Line 303-304: Beta-catenin is thought to exchange at beta-catenin degradasomes; this is clear from previous FRAP assays and the observation that phospho-beta-catenin accumulates in degradasomes upon proteasome inhibition (10.1158/1541-7786.MCR-15-0125). However, degradasome size hasn't, to my knowledge, been related to activity. Can this be clarified, please?

      We apologize for confusing β-catenin phosphorylation with β-catenin abundance. Here, we refer the catalytic activity of the DC to as the ability of the DC to promote β-catenin degradation rather than the kinetics of β-catenin phosphorylation. It is commonly observed that AXIN stabilization by TNKS inhibitors increases the DC size and reduces the β-catenin levels. As such, the induction of AXIN puncta by TNKS inhibitors is frequently used as an indicator of WNT/β-catenin pathway inhibition. However, we have found that, TNKS inhibition drives TNKS accumulation, which reduces the ability of the DC to promote β-catenin degradation. We agree that the DC only primes β-catenin but does not catalyze its degradation. We have revised our manuscript as follows: "increasing the local concentration of the DC components improves its 'effective activity'[50,51]."

      (10) There are previous hypotheses/proposals that the sensitivity of CRC cells to tankyrase inhibition correlates with APC truncation or PIK3CA status (10.1158/1535-7163.MCT-16-0578; 10.1038/s41416-023-02484-8). Have the authors considered expanding their cell line panel (Figure S7) to sample a wider range of cell lines, including some that are wild-type with regard to APC or Wnt/beta-catenin signalling in general? This would be a valuable addition to the work. Quantitated colony formation data could be moved to the main body of the manuscript.

      We have so far tested the effects of IWR1-POMA on the proliferation of DLD-1, SW480, HT-29, HCT116, and RKO cells (Fig. 6A and 6B). While a heterozygous Ser45 deletion in CTNNB1 confers resistance to IWR1-POMA, we did not observe sensitivity associated with APC or PIK3CA status. The ability of IWR1-POMA to suppress the growth of RKO cells expressing wild-type APC is consistent with a previous report that knockdown of both TNKS1 and TNKS2 stabilized PTEN to suppress cell proliferation and glycolysis in vitro and tumor growth in vivo (PMID: 25547115, Ref. 48) independently of the β-catenin pathway. We have added this new information as well as quantification of the colony growth results (Fig. S8A, S8B, S9A, S9F, and S9G) to the revised manuscript.

      (11) The manuscript only mentions toxicity (i.e., therapeutic window) in the last sentence of the Discussion section. As this is THE main challenge with tankyrase inhibitors (as mentioned above), can the authors expand their discussion of this aspect? Is there an expectation that PROTACs may be less toxic?

      As discussed above, evidence for on-target toxicity of WNT/β-catenin inhibition is mixed. Yet, the absence of dose-limiting toxicity for basroparib at doses up to 360 mg QD in human (PMID: 40964966, Ref. 64) is encouraging. PROTAC works by catalyzing target degradation, which is different from traditional catalytic inhibitors that require continuous target occupancy at a high level. It remains unclear whether the observed on-target toxicity of TNKSi is associated with TNKS accumulation at high doses, akin to the cytotoxicity induced by PARP1-trapping upon catalytic inhibition. We have included a brief discussion of the toxicity issue in the final paragraph of the Discussion section.

      (12) Figures 3, 4, 5A: For fluorescence microscopy experiments, can these be quantified, and can repeat data be included?

      We have included quantification data and replicate information for Fig. 3–5.

      (13) Figure 4, S6: An additional channel illustrating the distribution of cells (e.g., nuclei, cytoskeleton, or membrane) would be helpful for orientation and context for the AXIN1 signal.

      We have included cell outlines or nuclear staining for Fig. 3, 4, S6, and S7.

      (14) How were cytosolic fractions of cells prepared to assess cytosolic beta-catenin levels? This detail is missing from the methods.

      We have updated the Methods section to include additional details on the preparation of the cytosolic fractions of cells.

      Reviewer #3 (Public review):

      In this manuscript, Wang et al employ a chemical biology approach to investigate the differences between the enzymatic and scaffolding roles of tankyrase during Wnt β-catenin signalling. It was previously established that, in addition to its enzymatic activity, tankyrase 1/2 also plays a scaffolding function within the destruction complex, a property conferred by SAM-domain-dependent polymerization (PMID: 27494558). It is also known that TNKS1/2 is an autoregulated protein and that its enzymatic inhibition leads to accumulation of total TNKS proteins and stabilization of Axin punctae (through the scaffolding function of TNKS1/2), leading to rigidification of the DC and decreased β-catenin turnover. The authors surmised that this could, in part, explain the limited efficacy of TNKS1/2 catalytic inhibition for the treatment of colorectal cancers. To test this hypothesis, they evaluated a series of PROTAC molecules promoting the degradation of TNKS1/2 to block both the catalytic and scaffolding activities. They show that IWR1-POMA (their most active molecule) promotes more efficient suppression of beta-catenin-mediated transcription and is more active in inhibiting colorectal cancer cell and CRC patient-derived organoids growth. Mechanistically, the authors used FRAP to demonstrate that catalytic inhibitors of TNKS led to a reduced dynamic assembly of the DC (rigidification), whereas IWR1-POMA did not affect the dynamics.

      Overall, this is an interesting study describing the design and development of a PROTAC for TNKS1/2 that could have increased efficacy where catalytic inhibitors have displayed limited activity. Knowing the importance of the scaffolding role of TNKS1/2 within the destruction complex, targeting both the catalytic and scaffolding roles certainly makes sense. The manuscript contains convincing evidence of the different mechanisms of the PROTAC vs catalytic inhibitors. Some additional efforts to quantify several of the experiments and to indicate the reproducibility and statistical analysis would strengthen the manuscript. Ultimately, it would have been great to evaluate the in vivo efficacy of IWR1-POMA in an in vivo CRC assay (APCmin mice or using PDX models); however, I realize that this is likely beyond the scope of this manuscript.

      We thank the Reviewer for the helpful suggestions.

      I have some recommendations listed below for consideration by the authors to strengthen their study:

      (1) The title is slightly misleading, as it is already known that the scaffolding function of TNKS is important within the DC. The authors should consider incorporating the PROTAC targeting aspect in the title (e.g., PROTAC-mediated targeting of tankyrase leads to increased inhibition of betacat signaling and CRC growth inhibition).

      We have modified the title accordingly to "Targeting tankyrase scaffolding in the β-catenin destruction complex by PROTAC overcomes the limitation of catalytic inhibitors in cancer."

      (2) The authors should comment in the manuscript on the bell-shaped curve obtained with treatment of cells with the PROTACs (Figure S2C). This likely indicates tittering of the targets within a bifunctional molecule with increasing concentration (and likely reveals the auto-inhibition conferred by the catalytic inhibition alone).

      As suggested by the Reviewer, the bell-shaped dose-response likely originated from the formation of non-productive binary protein-ligand complexes at high PROTAC concentrations. We have added a sentence to clarify this unique behavior of PROTAC molecules.

      (3) The authors comment that using G007-LK as warehead was unsuccessful, but they do not show data. Do the authors know why this was the case?

      The structure-activity relationship of PROTACs is often unpredictable, as both the kinetics and thermodynamics of target and E3 ligase binding play important roles in promoting efficient target degradation. We have include data on G007-LK based PROTACs (Fig. S2D) in the revised manuscript.

      (4) Throughout the manuscript, the authors need to do a better job at quantifying their results (i.e., the western blots and the IF). For example, the degradation of TNKS1/2 in Figure 1D is not overly convincing. Similarly, the IF data in Figure 3 needs to be quantified in some ways. Along the same lines, the effect of IWR1-POMA treatments on the proliferation of cells and organoids should be quantified using viability assays... There is also no indication of how many times these experiments were performed and whether the blots shown are representative experiments. The quantification should include all experiments.

      We have included quantification of the immunofluorescence images, colony formation data, and Western blots in the revised manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) For clarity, can the authors use the official gene names, TNKS and TNKS2?

      We favor using TNKS1 and TNKS2 when referring to the protein for clarity and use TNKS for simplicity when referring to both proteins.

      (2) Line 92: The authors refer to TNKS2 "induction" - it remains unclear what is meant by "induction".

      We have changed "without induction" to "under basal conditions".

      (3) Can the authors please display molecular weight markers for Western blots throughout?

      (4) Line 144: The description "significantly more effectively" refers to Figure S5A, which shows a single, non-quantified Western blot. I don't think significance has been tested, and this statement should be reworded, or quantified aggregate data provided.

      We have added a Supplementary Information file showing molecular weight markers and quantification of Western blots.

      (5) Line 226: "plateaued at a much lower level" - can this be expressed more quantitatively in the text?

      We have included more quantitative information on the FRAP results.

      (6) Line 249: Can the authors repeat the cross-reference to Figure S7A here?

      We have repeated the cross-reference to the figures.

      (7) Line 266: The description of the experiment using the GSPT1/2 degrader CC-90009 would benefit from a brief recap of the purpose as not every reader will be familiar with this common PROTAC off-target. This is a very thorough analysis, though, and commendable.

      We have added background information on GSPT1 degradation to the revised manuscript.

      (8) Figure 1A: Can the number of repeats and the type of repeats be indicated, please?

      (9) Figure 2: Does n refer to biological or technical repeats?

      (10) Figure 5B, D: How many separate experiments are the data based on?

      (12) Figure S3D, S9A, D: number and types of repeats and the nature of the displayed data and error bars need to be included, please.

      (13) Figure S6B, S7B: I can see three data points, but it would still be helpful to state the number and type of repeats in the legend.

      (14) Figures S9A, S9D: There is value in showing the cumulative data from several repeats in the main figure (Figure 6, which currently is only qualitative) rather than the supplementary material.

      (15) Where single Western blots are shown, can the authors indicate how many experiments they are representative of?

      We have included the number of biological repeats for all data.

      (11) Figure S2C: For most graphs, the main response of interest occurs at low compound concentrations. The y-axis scale does not always help the reader to appreciate the effects, as the response seems small against the magnitude of the hook effect. Interrupting the y-axis as in the final panel may help, with y-axis scales consistent over all panels in the figure.

      We have updated Fig. S2C to emphasize on the degradation efficacy.

      (16) The authors may want to give further method details for some of their assays to facilitate replication of their experiments in the future. For example, the STF assay description is currently quite minimalistic. I assume the assay is fairly robust, though. Other details include cell media (general media details and specific additives and their concentrations in the 3D spheroid formation assay), etc. A general look at the methods section will likely be beneficial.

      We have updated the Methods section to provide more detailed experimental information.

      Reviewer #3 (Recommendations for the authors):

      (1) In Figure 2A, one of the most important findings of the manuscript is that IWR1-POMA induced promoted deeper suppression of beta-catenin-mediated transcription. This seems to be the case only at 3.2uM. Is it statistically significant? What are the data points on this graph? What are the error bars?

      We have included statistical analysis as Fig. S5G.

      (2) On Figure 2C and 2D, do the authors know why the TNKS20M1054V mutant is much better at promoting signaling than the TNKS1-PD ? Is it expression levels?

      It is indeed interesting that TNKS2-M1054V promoted significantly stronger WNT signaling than TNKS1-PD. The basis for its strong scaffolding effect is unclear.

      (3) In Figure 4C, the authors claim that when cells are treated with IWR1-POMA, AXIN1 is distributed diffusely throughout the cytoplasm. It appears that small punctae are visible.

      Quantitative analysis (Fig. 4F) suggest that the size of AXIN1 puncta upon IWR1-POMA is rather insignificant.

      (4) Label on Figure 1D has a spelling error TNKS1/2.

      Corrected.

    1. eLife Assessment

      This manuscript presents openretina, an open-source platform that integrates retinal datasets, model training, benchmarking, and in silico analysis tools within a unified framework. The resource is valuable because it addresses long-standing challenges in reproducibility, accessibility, and cross-study comparison in computational retina research, while providing a foundation for community-driven model development and evaluation. The supporting evidence is solid, with the authors demonstrating a functional and well-documented platform across multiple datasets and species, although a clearer discussion of model interpretability, current performance limitations, and data quality standards would strengthen the resource.

    2. Reviewer #1 (Public review):

      Summary:

      This "Tools and Resources" submission describes a platform for the modeling of stimulus-response relationships in the retina. It includes a repository for experimental data sets with standardized programmatic access, and a suite of software for constructing stimulus-response models and evaluating them.

      Strengths:

      (1) The paper is well written.

      (2) The platform could serve an integrative function by connecting different research programs and offering a common baseline for evaluating stimulus-response models.

      (3) The finding that there is "substantial explainable variance remains uncaptured by current models" is a useful insight to motivate further work and measure progress.

      Weaknesses:

      (1) The modeling supported by the package focuses on predictive accuracy at the cost of less interpretability.

      (2) The article needs to make a stronger argument that this style of modeling is fruitful, especially when applied to the retina.

      Main comments:

      (1) Abstract machine learning vs mechanistic models. The "Core + Readout" architecture advocated here seems to be divorced from all the neurobiological detail that is already known in the retina. It mostly aims at prediction, not interpretation. Such a black-box modeling framework is useful in brain regions where we know very little about connectivity, or mechanisms, or even about the primary function being performed there, like in the mammalian cortex. In those cases, any model that can deliver a prediction is a step forward, even if it does not connect to biological mechanisms. But that's decidedly not the situation in the retina, where so many mechanistic details are known: from consensus cell types, to synaptic detail, to single-neuron biophysics, to circuit motifs. How can one connect this ML modeling approach with the extensive mechanistic knowledge available in retinal neuroscience? And can the combination somehow lead to a better understanding? The authors seem to recognize this tension (e.g. line 215ff and 370ff) but don't give it much weight. A stronger case needs to be made here for how this kind of modeling will advance the field.

      (2) The "gradient field" approach. Figure 4c illustrates a case of this dissonance. The gradient field of the response increases with contrast in multiple directions. This is obvious a priori (see line 274) from the more mechanistic model we already have of this On-Off cell. These are the W3 cells described in www.pnas.org/cgi/doi/10.1073/pnas.1211547109. The circuit-based model from that paper, with rectifying on and off subunits from bipolar cells, gives a much more compact explanation for what the neuron does. Because each of the subunits has a spatio-temporal receptive field, this model can predict the entire dynamics to arbitrary stimuli, rather than just 2 dimensions of static stimuli as in the present analysis. So what is the value added here? Again, a stronger case needs to be made that these "Core + Readout" modeling activities enhance understanding.

      (3) The "most exciting input" approach (Line 193ff):

      - Presumably, some power constraint must be put on the stimulus? Otherwise, increasing the contrast will make it more exciting. What are these constraints?

      - Presumably, this optimal stimulus is computed from the model based on non-optimal stimuli? What are the assumptions going into that?

      - The most exciting stimulus is not necessarily the most useful characterization. Near its maximal firing rate, the neuron doesn't discriminate stimuli much, because the slope there is zero (line 237). Instead (or in addition), one would like to know along which stimulus axis the neuron is most sensitive. See e.g. discussion in Dayan & Abbott 2000, Figure 3.11.

    3. Reviewer #2 (Public review):

      Summary

      openretina is a Python package for training and applying convolutional neural network-based models of retinal ganglion cell responses. The package integrates dataloading, model training, and evaluation in a unified framework built on PyTorch Lightning and Hydra, and ships with pre-trained model checkpoints and publicly available datasets (whitenoise, natural scenes) spanning multiple species (marmoset, mouse, axolotl, salamander) and recording modalities (multielectrode array recordings or 2-p calcium imaging). Beyond predictive modelling, openretina includes a suite of in silico analysis tools for probing learned representations, including maximally exciting input synthesis, discriminatory stimulus optimisation, and model weight visualisation. The broader openretina initiative aims to establish a community-driven platform for computational retina research, lowering barriers to entry and facilitating cross-dataset model benchmarking. This is a valuable contribution given the longstanding fragmentation of datasets, codebases, and analysis practices across retina laboratories.

      Strengths:

      The tool has several strengths. By providing a framework built on deep learning infrastructure, the package substantially lowers the barrier to entry for researchers without extensive machine learning backgrounds. The inclusion of pre-trained model checkpoints across multiple species and recording modalities will allow users to apply state-of-the-art models. The in silico toolkit - and in particular the MEI synthesis pipeline - has already demonstrated its scientific potential, with prior work using optimised stimuli to discover a previously uncharacterised RGC type confirmed experimentally, illustrating what becomes possible when these tools are made broadly accessible. The current modular Core + Readout architecture is a well-suited architecture for modeling retina responses. The HDF5-based data standard provides a sensible common format for contributing new datasets. Overall, the initiative is well-motivated, the engineering is competent, and the vision of a collaborative, community-driven platform for retina modelling is one that the retina community would benefit from.

      Weaknesses:

      (1) The in silico tools provided are valuable, but users should interpret their outputs in light of the performance of the underlying models. The predictive performances of current models and datasets in the package are far from performance ceilings.

      (2) The authors appropriately note that optimised stimuli reveal what a neuron responds to but not how the computation is implemented. I would encourage readers to keep this distinction in mind when using the weight visualization tools as well - convolutional filters in a shared, unconstrained core do not map onto retinal circuit elements, and should be treated as model descriptors rather than circuit proxies. For example, RGCs of the same type may appear to sample inputs from two different filters, which should have been a single filter. Or a single RGC may be sampling from two filters, which under more constrained conditions could be approximated with a single filter. These are degeneracies in the CNN modeling framework that should be kept in mind when drawing circuit-level interpretations.

      (3) The datasets currently distributed with the package vary in recording quality, and users should be aware that model performance may not only reflect architectural limitations but may also be limited by noise and data artifacts, including spike sorting errors.

      (4) As the platform grows and community-contributed datasets are added, explicit data quality standards will be essential. I encourage the authors to develop dataset standards to ensure that their resource provides access to highly curated datasets, which I believe is an important step in having high-fidelity models whose functional interpretations can be trusted.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript presents openretina, a Python-based platform designed to facilitate collaborative retinal modeling across datasets, laboratories, species, and recording modalities. The package provides standardized model architectures, evaluation metrics, and analysis tools, while also integrating several publicly available retinal datasets. The authors further demonstrate the platform through examples of in silico analyses and model benchmarking.

      Strengths:

      (1) Emphasis on standardization and reproducibility. Retinal modeling has become increasingly dependent on deep learning approaches, yet datasets and evaluation procedures remain fragmented across laboratories. By providing a unified framework, the authors lower barriers to entry and create opportunities for more systematic comparisons of models and datasets.

      (2) The manuscript is clearly written, and the examples effectively illustrate the range of analyses supported by the platform.

      (3) The benchmarking results are useful, particularly because they reveal substantial remaining gaps between current model performance and explainable variance ceilings.

      Weaknesses:

      Not a weakness per se, but rather a limitation, is that the manuscript focuses on software infrastructure rather than new biological or computational insights. While this is appropriate for a resource paper, some of the scientific examples, such as the gradient-field analysis of ON-OFF cells, function more as demonstrations than as rigorous validations of novel hypotheses. It might be useful to add a few sentences discussing potential scientific projects that can be immediately facilitated by the openretina (the current text in the Discussion focuses more on advancements in the technical/social aspects of science that will be supported by openretina).

      Overall, this is a valuable and timely resource that is likely to benefit the retinal and computational neuroscience communities.

    5. Author response:

      We thank the editors and reviewers for their thoughtful assessment of our manuscript, and for recognizing openretina as a valuable and timely resource for the retinal modelling community.

      We are especially glad that the reviewers appreciated the motivation of the project, the focus on standardization and reproducibility, and the potential of the platform to support systematic benchmarking and community-driven model development.

      We also understand the concerns raised. In the revision of the manuscript, we will strengthen the conceptual discussion of how predictive models, including the current “Core + Readout” models, can contribute to retinal neuroscience alongside more mechanistic and circuit-based approaches. This is a central matter for us, and one that some of us have recently addressed in a broader review on current trends in retina modelling (see https://doi.org/10.1016/j.visres.2026.108854). We will draw on this perspective to better articulate when predictive models are useful, where their limitations lie, and how openretina can provide infrastructure for comparing functional, normative and mechanistic models within a shared framework.

      We will also clarify the scope and limitations of the in-silico analysis methods provided within openretina. This will include a more explicit discussion of how MEIs, gradient-field analyses, and model-weight visualisations should be interpreted.

      Furthermore, we will add more information that will help the reader better judge different aspects of dataset quality, including, for example, spike-sorting or calcium-processing information and explainable-variance distributions. We note, however, that there are many subtle details about experimental workflows that are difficult to capture in compact indicators. In addition, we will make it clearer that the manuscript represents a snapshot of a living resource: The website, dataset cards, documentation, and repository will be the primary source of this information, especially as new datasets are contributed.

      Finally, we will of course address the technical clarifications raised by the reviewers, with the aim of making the manuscript more accessible overall.

      We are grateful for the reviewers’ constructive comments and believe that addressing these points will make our presentation of openretina clearer and more useful to the community.

    1. eLife Assessment

      In this manuscript, the authors describe a cell-specific mechanism by which glutamate transporters regulate the fidelity with which T-stellate cells in the mouse ventral cochlear nucleus relay information from auditory nerve inputs. The study is supported by solid electrophysiological data. It provides valuable insights into how the rapid binding of glutamate to transporters shapes auditory information processing at specific synapses.

    2. Reviewer #1 (Public review):

      In this article, the authors investigate how glutamate transporter function regulates excitability and synaptic coding in T-stellate cells in the mouse ventral cochlear nucleus. They test this in acute brain slices using whole-cell electrophysiology and artificially raise the relative local concentration of glutamate via pharmacological inhibition of transporter proteins. The main finding is that when sub-saturating doses of DL-TBOA are applied, cells become much more sensitive to synaptic input, diminishing the normally high fidelity of EPSP-spike coupling in these neurons. Notably, high-frequency stimulation in the presence of DL-TBOA reveals a large and slowly decaying AMPA receptor component that underlies persistent/rebound firing in earlier recordings. These effects are not seen in other ventral cochlear neurons, suggesting that rapid glutamate clearance in T-stellate cells, particularly, is important for auditory intensity coding. Overall, these experiments are well-performed, and the findings are robust, though there are some aspects that could be expanded to make the work more impactful. These include a better understanding of the relative contribution of neuronal vs glial transporters and an ability to separate the relative contributions of tonic glutamate concentrations in the cleft vs changes in membrane potential in action potential output. Additionally, there were some minor issues of clarity in both the figure presentation and the main text language that should be addressed.

      Major Points:

      (1) Given the dramatic effect of saturating DL-TBOA on tonic leak/RMP and that the sub-maximal concentration used in most of the experiments still varied between 25-50 uM, Figure 1 would be strengthened substantially by a dose-response curve. Ideally, 5 or 6 concentrations, plotting the effect on tonic current or RMP increase.

      (2) Examining the contribution of glial (EAAT1/2) vs. neuronal (EAAT3) transporters (Fig 8) is intriguing but comes across as incomplete here, especially given the small number of recordings. Using a different non-selective EAAT inhibitor (TFB-TBOA) to chase the EAAT1/2 blocker combo seems like an odd choice, given that you have already characterized the effects of DL-TBOA well. One could also try a lower concentration (~50-100 nM) of TFB-TBOA since it is somewhat selective itself for glial EAAT1/2. Given the data presented, neuronal transporters (presumably EAAT3) appear to dominate the rapid clearance of glutamate at this synapse, but this point isn't emphasized or explored sufficiently.

      (3) Separating the effects of depolarization vs. glutamate clearance was never explored. What effect does depolarizing the cell ~10 mV in control conditions (i.e., without TBOA) have on AP number/fidelity during synaptic stimulation experiments? The authors state that submaximal DL-TBOA generally causes no more than a 5 mV change in RMP, but tonic depolarization could also influence spike fidelity. This experiment could demonstrate that the increase in excitability during/after stimulation is not due to increased engagement of voltage-gated channels.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript addresses an important and mechanistically interesting question: whether plasma membrane glutamate transporters contribute only to slow clearance of ambient glutamate or whether they can rapidly shape synaptic signaling during high-frequency auditory activity. This manuscript provides important evidence that EAAT-mediated glutamate uptake is not merely a slow background clearance mechanism but is essential for maintaining reliable synaptic transmission and linear stimulus-intensity coding in ventral cochlear nucleus T-stellate cells during sustained auditory nerve activity.

      Strengths:

      The finding that EAATs may be required for rapid, local control of glutamate during high-frequency auditory nerve activity is interesting and could have broad relevance to auditory processing. The electrophysiological evidence is generally strong, particularly the use of patch-clamp recordings, stimulus trains, partial versus complete EAAT blockade, and comparison with bushy cell/endbulb synapses. The comparison between T-stellate cells and bushy cells/endbulb synapses strengthens the manuscript. The authors demonstrate that EAAT blockade disrupts coding in T-stellate cells but has little effect on bushy cell spike transmission, supporting a cell-type- and synapse-specific role of glutamate uptake.

      Weaknesses:

      However, some mechanistic conclusions, especially the specific contribution of neuronal versus glial EAATs and the absence of glutamate crosstalk between auditory nerve inputs, rely mainly on pharmacological and indirect electrophysiological inference and would be strengthened by additional anatomical, genetic, or direct glutamate-sensing evidence.

      (1) Clarification of DL-TBOA concentration.

      The authors used bath application of 200 µM TBOA and 25-50 µM in the other experiments, stating that "sub-maximal concentrations (25-50 µM)". The authors should provide a clearer rationale for why different concentrations were used across experiments rather than a fixed concentration.

      The reversibility of DL-TBOA effects should be demonstrated by washout experiments. In addition, potential off-target effects of DL-TBOA on postsynaptic receptors, intrinsic membrane excitability, or presynaptic release (e.g., PPR measurement) should be carefully considered. It would also be useful to test the effects of the submaximal DL-TBOA concentrations (25-50 µM) on membrane potential and inward currents, shown in Figure 1, to determine whether these concentrations depolarize the membrane potential in current-clamp mode or induce inward currents under voltage-clamp conditions.

      (2) Potential contribution of altered intrinsic excitability.

      In Figures 3B and 3C, DL-TBOA appears to induce additional action potentials even immediately after the first stimulation, whereas Figures 6 and 7 suggest that the first EPSC is not substantially altered. This raises the possibility that the enhanced firing may partly result from a modest depolarization caused by background glutamate accumulation or from other changes in intrinsic membrane properties after drug treatment. To address this, the authors should provide a quantitative analysis of physiological parameters under submaximal DL-TBOA conditions, including spontaneous action potential frequency, resting membrane potential, input resistance, and spike threshold.

      (3) Spillover/ crosstalk between AN-fiber-synpases.

      The authors should provide more explanation of how altering the number of active auditory nerve fibers demonstrates the absence of glutamate spillover/crosstalk between bouton synapses. Strong stimulation likely recruits more AN fibers, but it may also change release probability, axonal synchrony, or stimulation spread. The authors should more clearly justify the interpretation that strong stimulation recruits additional independent AN fibers rather than altering release probability or activating fibers with different intrinsic properties.

      (4) Interpretation of glial versus neuronal EAAT contributions.

      The authors claim that both neuronal and glial transporters contribute to rapid uptake using pharmacological approaches. The pharmacological data demonstrate that glial EAATs play a major role in glutamate clearance at T-stellate cell synapses. The strong increase in EPSC decay time and synaptic charge after UCPH-101/DHK application supports the conclusion that glial transporters contribute substantially to limiting glutamate accumulation during sustained auditory nerve activity. However, the conclusion that neuronal EAATs contribute directly should be stated with some caution. The evidence for neuronal EAAT involvement is indirect and depends on the pharmacological specificity and completeness of glial EAAT blockade. The conclusion would be strengthened by additional evidence, such as EAAT subtype expression/localization in T-stellate cells or auditory nerve terminals, transporter current recordings, immunohistochemistry, or genetic manipulation of neuronal EAATs. In addition, fitting the decay phase with a double-exponential model may help determine whether glial and neuronal EAATs contribute over distinct temporal windows.

    1. eLife Assessment

      This manuscript describes an important development of several variants of optogenetic tools to control endogenous p53 activity. They are based on peptides competing with Mdm2/MdmX for binding to p53, thus releasing p53 from its negative regulators and stabilizing its cellular levels. In principle, the data are convincing but should be complemented by investigations of p53 target genes at endogenous levels (instead of only reporter constructs). The study therefore remains incomplete but will be of interest to scientists working on optogenetics as well as the p53 field.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, the authors apply the AsLOV2 domain to control the localisation and the exposure of two peptides (PMI and PMI-M3) that compete with Mdm2/MdmX for binding to p53, thus freeing p53 from these negative regulators and allowing its levels to rise. The authors follow an established strategy in optogenetics, which is to combine two layers of regulation for tighter control: (1) caging the peptide into the Ja helix of AsLOV2; 2) sequestration of the peptide away from its site of action using the LOVTRAP system.

      Strengths:

      The authors show that a reporter is activated when cells are exposed to light. A strength is in the lower background that was achieved after adding the second layer of regulation.

      Weaknesses:

      This study claims to be focused on the control of endogenous p53; however, endogenous p53 levels are not quantified. Moreover, endogenous p53 target genes are also not analysed. Only a synthetic reporter is quantified, which has been placed in the genome of HCT116 cells after the creation of a stable cell line. Microscopy images show only one or a maximum of two cells. Finally, the authors claim their strategy is a general one that can be applied to control other peptides, but they do not show this generality in this paper.

    3. Reviewer #2 (Public review):

      The authors developed Opto-MDMi, an optogenetic system for light-controlled activation of endogenous p53. The main idea is to target the p53-MDM2/MDMX regulatory interaction using PMI inhibitory peptides. This is a nice strategy because it avoids overexpression of p53, which can have adverse effects that might confound the study of p53 activity. The authors first tested a LOVTRAP-based localization strategy, which showed some efficacy but also showed basal activation. They then developed a LOV2-PMI peptide-caging module to control the activity of the PMI peptide itself, testing for interactions first in vitro and then in vivo. Finally, they combined the two systems into a dual-lock design, where LOVTRAP controls localization and LOV2-PMI controls peptide activity. This combination led to somewhat more potent stimulation of p53 activity.

      Another useful aspect of the paper is the detailed description of the development and testing of the LOV2-PMI peptide-caging module, which may aid in the design of other LOV2-based peptide-caging designs.

      Strengths

      Overall, the paper is novel and rigorous, and the claims are supported by the data. The optoMDMi tool seems ready for implementation, for example, to manipulate and study the role of p53 signaling dynamics. A few points of clarification would strengthen the work.

      Weaknesses

      The authors develop many tool variants, but there is some lack of clarity over how all of these tools compare to each other, and which ones interested users should use. The work would also be strengthened by showing modulation of endogenous p53 in more than one cell line.

    1. eLife Assessment

      Verma and colleagues interrogate the mechanisms of phagosome maturation arrest during Mycobacterium tuberculosis infection. While cellular events that culminate in this arrest have been largely elucidated, involvement of other organelles, such as mitochondria, has not been highlighted mechanistically. In this valuable study, elements of mitochondrial quality control, such as mitophagy and mitochondrial-derived vesicles involvement, are shown to be paramount in the host-pathogen tussle. The evidence supporting the main conclusions is solid, based on multiple complementary approaches and appropriate controls, although some central mechanistic aspects of the proposed pathway remain only partially resolved.

    2. Reviewer #1 (Public review):

      Summary:

      This is an important and interesting manuscript that uncovers the cross-talk between mitochondrial quality control and phagosome maturation arrest imposed by Mtb.

      A broader host pathogen (intracellular) question pertains to evading phagosomal maturation/arrest. While cellular events that culminate in this arrest have been largely elucidated, involvement of other organelles, such as mitochondria, has not been highlighted mechanistically. This manuscript paints a larger picture than the well-known conventional endolysosomal pathway and portrays a larger landscape involving elements of the mitochondrial quality control, such as mitophagy and mitochondrial-derived vesicles' involvement in the host-pathogen tussle.

      Strengths:

      The systematic characterisation to unravel the interplay between mitochondrial-related pathways and the endolysosomal system allows the authors to unearth some important findings.

      Weaknesses:

      The conclusions drawn require more robust experimentation and analysis.

    3. Reviewer #2 (Public review):

      This manuscript examines the role of autophagy receptor proteins, particularly p62/SQSTM1, in regulating intracellular Mtb survival in human macrophages. Counterintuitively, depleting p62 reduces bacterial survival rather than enhancing it, pointing to a previously unrecognised mechanism. The authors demonstrate that in the absence of p62, mitochondrial quality is maintained through enhanced TOM20⁺ mitochondria-derived vesicle (MDV) biogenesis, dependent on MIRO1/MIRO2. During Mtb infection, these MDVs are redirected to bacterial phagosomes, promoting RAB7 recruitment, overcoming phagosome maturation arrest and facilitating lysosomal targeting of Mtb. In parallel, bacteria experience increased oxidative stress, further contributing to bacterial killing.

      Strengths:

      The mechanistic chain is built using multiple complementary approaches, including genetic perturbation, redox biosensors, metabolic assays and microscopy. The use of primary human macrophages from multiple donors alongside established cell lines increases confidence that the phenotype is not cell-line specific. The replication clock experiment is particularly elegant and clearly demonstrates that the reduction in bacterial burden reflects enhanced killing rather than impaired bacterial replication. Overall, the study identifies an unexpected connection between mitochondrial quality control and phagosome maturation and provides a potentially important advance in our understanding of host-pathogen interactions.

      Weaknesses:

      The study remains entirely in vitro, and the phenotype is absent in mouse macrophages, limiting the immediate physiological and translational relevance of the findings. In addition, many of the central mechanistic conclusions rely heavily on colocalisation analyses, making it difficult to distinguish direct mechanistic relationships from associated trafficking events.

      Overall, this is an interesting and technically strong study that uncovers a novel link between mitochondrial quality control and anti-mycobacterial defence. The mechanistic model is plausible and supported by substantial experimental work. However, several aspects of the proposed pathway require stronger experimental support before some of the broader conclusions can be fully justified.

      Major points

      (1) The central conclusion that TOM20⁺ MDVs are recruited to Mtb-containing phagosomes is based largely on microscopy and colocalisation analyses. Additional orthogonal approaches would strengthen this key aspect of the study and help establish the nature of the vesicles recruited to bacterial phagosomes.

      (2) The proposed mechanism whereby TOM20⁺ MDVs facilitate RAB7 recruitment and reverse phagosome maturation arrest remains incompletely demonstrated. While the MIRO1/2 and RAB7 knockdown experiments support the model, they do not directly establish a causal link between MDV recruitment and phagosomal RAB7 acquisition. Additional experiments addressing this step would considerably strengthen the manuscript.

      (3) The absence of a phenotype in mouse macrophages raises important questions regarding the conservation and physiological relevance of the proposed mechanism. The authors should discuss possible explanations for this species-specific effect and, if feasible, provide additional experimental insight into the basis of this difference.

      (4) The conclusion that mitochondrial quality is maintained despite impaired p62-dependent mitochondrial turnover is based primarily on mitochondrial content, membrane potential, ROS measurements and Seahorse analysis. These are informative but relatively indirect measurements. Additional assessment of mitochondrial turnover by mitophagy would strengthen this aspect of the study.

      (5) The proteins studied throughout the manuscript (p62/SQSTM1, NDP52, OPTN, TAX1BP1 and NBR1) are generally classified as selective autophagy receptors rather than adaptors. The terminology should be corrected throughout the manuscript.

      Minor points:

      (1) Several conclusions throughout the manuscript are based primarily on colocalisation analyses. The limitations of these approaches should be acknowledged explicitly.

      (2) The discussion would benefit from a clearer consideration of how the proposed mechanism relates to established pathways regulating phagosome maturation arrest during Mtb infection.

      (3) The authors may wish to comment on whether enhanced MDV biogenesis could represent a broader host defence mechanism against intracellular pathogens beyond Mtb.

    1. eLife Assessment

      This valuable study provides a cross-species single-cell transcriptomic resource for early female gonadal development in mammals. The data supporting the main conclusion remain incomplete, and experimental validation is needed to strengthen the conclusions. The work will be of interest to reproductive biologists and developmental biologists.

    2. Reviewer #1 (Public review):

      Summary:

      Fang et al. characterize the cellular basis of early ovarian development through a comparative analysis of single-cell transcriptomic data. The authors integrate a novel bovine scRNA-seq dataset, spanning six gestational stages (E38-E112), with stage-matched human (PCW6-16) and mouse (E11.5-E18.5) counterparts. Beyond identifying shared gonadal cell types across these three species, the study uncovers a previously uncharacterized bovine-specific cell population with steroidogenic features. Their analysis highlights conserved, dynamically expressed regulators, including TFAP2C and ZCWPW1 in germ cells and FOS and JUNB in granulosa cells. Furthermore, by employing a machine learning Support Vector Machine (SVM) model, the authors quantify cell-type conservation, demonstrating that while immune and germ cells are highly conserved across species, granulosa cells exhibit substantial evolutionary divergence. This study makes a significant contribution to developmental biology by establishing a comprehensive, cross-species single-cell roadmap of fetal ovarian development. By integrating livestock data with human and rodent models, the authors identify novel cellular states and provide a framework for assessing transcriptional conservation across species.

      Strengths:

      (1) While human and mouse fetal ovaries have been mapped, the inclusion of a high-resolution bovine dataset (107,930 cells total across the study) provides a critical "large mammal" perspective that is often missing from comparative studies.

      (2) The identification of a bovine-specific cell population is an important finding. It suggests that ruminants may have a different developmental timeline for steroidogenic precursors (potentially theca cell ancestors) compared to rodents or humans.

      (3) Training a Support Vector Machine (SVM) to quantitatively assess cell-type similarity is a major strength. It moves beyond qualitative UMAP "eye-balling" to provide a statistical probability of conservation.

      (4) The study links gene expression to higher-order biological processes like epigenetic reprogramming and cell-cell communication (CellChat), providing a holistic view of the gonadal niche.

      Weaknesses:

      (1) The authors integrated publicly available scRNA-seq datasets generated across different laboratories and technical platforms. However, the specific methods used to control for and evaluate batch effects are not clearly described. It is critical to clarify whether the observed species-specific differences are purely biological or partly influenced by technical variation between datasets.

      (2) A challenge inherent to all single-cell studies is the reliance on manual marker-gene-based annotation. While this is standard practice, it remains unclear how robust these assignments are, particularly for the novel "bovine-specific" population. Further evidence or cross-validation (e.g., through varied clustering resolutions or automated annotation tools) is required to ensure these clusters represent true biological states rather than computational artifacts.

      (3) The authors utilized a linear SVM to assess cross-species similarity. However, it is not clear how this model performs compared to established single-cell mapping and comparative tools (e.g., MetaNeighbor or Seurat v5). Providing a justification for this specific SVM-based approach, or a brief comparison with existing benchmarks, would strengthen the methodological rigor of the study.

      (4) While the computational evidence is compelling, the study would be significantly enhanced by independent validation of the "unclassified bovine-specific" cell population. To confirm the biological reality and reproducibility of this novel cell state, the authors should provide additional evidence. This could include in situ validation (e.g., immunofluorescence or in situ hybridization) to determine its physical location and morphology within the gonad, or demonstrating the presence of this specific cell population within an independent, non-overlapping bovine dataset.

    3. Reviewer #2 (Public review):

      Summary:

      The authors generate a comparative single-cell transcriptomic atlas of fetal ovarian development in cattle, human, and mouse, with the goal of identifying conserved and species-specific cellular and molecular features of early ovarian differentiation. The study provides a valuable resource for the field and reveals potentially interesting species-specific characteristics, including a putative bovine steroidogenic cell population. While the dataset is substantial and the computational analyses are generally appropriate, several major conclusions rely primarily on computational inference without independent experimental validation, limiting the strength of evidence supporting some of the central claims.

      Strengths:

      This study provides a valuable cross-species single-cell atlas of fetal ovarian development by integrating newly generated bovine data with human and mouse datasets. The work fills an important gap in reproductive biology and offers a useful resource for investigating conserved and species-specific features of ovarian development.

      The analyses are comprehensive and combine developmental trajectory reconstruction, regulatory network inference, cell-cell communication analysis, and cross-species classification. The identification of a putative bovine-specific steroidogenic cell population is particularly intriguing and may provide a basis for future studies of species-specific ovarian development.

      Weaknesses:

      The main limitation is that several key conclusions rely primarily on computational analyses without independent experimental validation. In particular, the proposed bovine-specific steroidogenic cell population, which represents the major novel finding of the study, is supported only by transcriptomic evidence.

      In addition, many mechanistic interpretations derived from trajectory, regulatory network, and cell-cell communication analyses remain speculative. While the study succeeds as a comparative resource, the evidence supporting several of the central biological claims remains incomplete, and the biological significance of some cross-species differences is not fully explored.

    1. eLife Assessment

      The Review Article by Bal and co-workers presents an overview of skeletal muscle physiology, focusing on Sarcoplasmic Reticulum and Mitochondria-Associated Membranes (MAMs) in relation to calcium handling, ROS, and signaling. It provides a foundation based on the current literature but could have gone further by identifying future research directions and potential avenues for therapeutic intervention.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript is a narrative review addressing age-related alterations in SR-mitochondria interactions in skeletal muscle and their contribution to sarcopenia. It synthesizes existing literature on calcium signaling, mitochondrial dynamics, redox balance, and structural remodeling, and discusses potential interventions including exercise and pharmacological strategies. While the topic is timely and relevant, the manuscript largely reiterates established concepts without providing sufficient conceptual novelty, critical synthesis, or mechanistic insight beyond the current literature.

      Strengths:

      (1) Timely topic: The focus on SR-mitochondria communication in aging muscle is relevant and of growing interest.

      (2) Broad coverage: The review compiles a wide range of literature spanning calcium handling, mitochondrial biology, ROS signaling, and exercise physiology.

      (3) Clear organization: The manuscript is structured logically with thematic sections (SR, mitochondria, MAMs, aging, interventions).

      (4) Didactic value: Could serve as a general overview for non-specialists entering the field.

      Weaknesses:

      (1) Lack of novelty and conceptual advance: The manuscript does not offer new hypotheses, frameworks, or critical reinterpretation of the field. Most statements summarize already well-established knowledge, and no unifying model or novel perspective is developed to justify publication in a high-impact journal like eLife.

      (2) Limited critical analysis: The review is predominantly descriptive rather than analytical. Conflicting findings (e.g., MFN2 roles, MAM density changes, Ca²⁺ overload vs deficiency) are mentioned but not critically evaluated or reconciled. There is little discussion of limitations in the cited studies or gaps in the field.

      (3) Overgeneralization and speculative claims: Several assertions are presented with insufficient nuance (e.g., causal links between MAM disruption and sarcopenia, or therapeutic efficacy of interventions). The distinction between correlation and causation is often unclear, reducing scientific rigor.

      (4) Insufficient depth for a specialist audience: Despite its length, the manuscript lacks mechanistic depth in key areas (e.g., precise molecular regulation of MAMs in vivo, tissue-specific differences, quantitative aspects of Ca²⁺ flux). It reads more like a textbook summary than a high-level scholarly review.

      (5) Redundancy and verbosity: Many sections repeat similar concepts (Ca²⁺ dysregulation, ROS effects, mitochondrial dysfunction) without adding new insight, leading to an unnecessarily long manuscript with limited added value.

      (6) Weak integration of recent literature into a coherent narrative: Although many references are cited, they are not effectively synthesized into a cohesive argument. The manuscript lacks a strong central thesis or clearly defined take-home messages.

      (7) Limited translational or experimental perspective: The section on therapeutic targeting is largely speculative and does not critically assess feasibility, limitations, or current clinical evidence.

    3. Reviewer #2 (Public review):

      This review addresses a highly relevant and timely topic, namely the role of sarcoplasmic reticulum-mitochondria communication and mitochondria-associated membranes (MAMs) in skeletal muscle aging. The manuscript successfully brings together literature from several interconnected fields, including calcium signaling, mitochondrial biology, excitation-contraction coupling, muscle metabolism, and sarcopenia. Given the growing interest in organelle crosstalk as a determinant of muscle health and disease, the topic is undoubtedly of considerable interest to the readership and has the potential to make a valuable contribution to the field.

      However, in its current form, the manuscript devotes a substantial proportion of its content to the description of well-established concepts that are already extensively covered in the literature. Large sections are dedicated to general skeletal muscle physiology, excitation-contraction coupling, calcium handling, mitochondrial biology, and MAM structure and composition. While this background information is useful, the level of detail is often excessive for a review that aims to focus on aging-induced alterations in SR-mitochondria interactions. As a consequence, the central theme of the manuscript becomes diluted, and the review reads more like a broad overview of skeletal muscle physiology than a focused analysis of aging-related MAM remodeling.

      In contrast, the sections specifically dedicated to aging and MAM dysfunction, which represent the most novel and potentially impactful aspects of the review, are comparatively brief and largely descriptive. The discussion of how aging alters MAM architecture, calcium microdomains, mitochondrial calcium signaling, and organelle communication would benefit from substantially greater depth. For example, although the manuscript highlights alterations in proteins such as MFN2, IP3R, VDAC, and MCU, the mechanistic implications of these changes for sarcopenia and age-associated muscle dysfunction are not critically developed. Similarly, the review would be strengthened by a more comprehensive discussion of the evidence linking MAM disruption to impaired muscle performance, metabolic inflexibility, denervation, and mitochondrial dysfunction during aging.

      Another limitation is that much of the manuscript summarizes published findings without sufficiently evaluating the strength of the evidence or discussing existing controversies. Several statements imply causal relationships between MAM disruption and sarcopenia, whereas in many cases, the available data remain largely correlative. The authors should more clearly distinguish between established mechanisms, experimental observations, and emerging hypotheses. A more critical assessment of conflicting findings, particularly regarding the role of MFN2 and the dual consequences of altered mitochondrial calcium uptake, would considerably improve the scientific rigor of the review.

      A major omission concerns the role of mitochondrial Ca²⁺ uptake in skeletal muscle physiology and aging. Throughout the manuscript, mitochondrial Ca²⁺ uptake is presented as a central determinant of muscle function and as a key mechanism linking MAM disruption to sarcopenia. However, the authors do not adequately discuss evidence that challenges this view. In particular, genetic mouse models lacking MCU exhibit surprisingly mild skeletal muscle phenotypes under basal conditions despite a near-complete abolition of rapid mitochondrial Ca²⁺ uptake. These findings have generated considerable debate regarding the physiological importance of mitochondrial Ca²⁺ uptake for muscle function and metabolic regulation. While MCU deletion clearly affects exercise adaptation and certain stress responses, the relatively modest baseline phenotype suggests the existence of compensatory pathways and raises important questions regarding the extent to which impaired mitochondrial Ca²⁺ uptake alone can explain age-associated muscle dysfunction. A balanced review should acknowledge these observations and discuss the ongoing debate regarding the relative contributions of mitochondrial Ca²⁺ deficiency versus mitochondrial Ca²⁺ overload in aging skeletal muscle.

      Similarly, the discussion of MFN2 would benefit from greater nuance. The manuscript largely presents MFN2 as a structural tether linking the sarcoplasmic reticulum and mitochondria. However, MFN2 is a multifunctional protein with well-established roles in mitochondrial fusion, mitochondrial network organization, mitophagy regulation, and metabolic signaling. Consequently, many of the phenotypes associated with altered MFN2 expression cannot be unequivocally attributed to changes in MAM formation. The review does not sufficiently distinguish between the effects of MFN2 on organelle tethering and its effects on mitochondrial dynamics. This distinction is particularly important because several studies have questioned whether MFN2 acts primarily as a positive tether, a negative regulator of contacts, or whether its influence on organelle communication is secondary to its role in controlling mitochondrial morphology. As a result, attributing age-related alterations in SR-mitochondria communication solely to changes in MFN2-mediated tethering may oversimplify a considerably more complex biological scenario.

      The manuscript's organization could also be improved. The sections discussing aging-related alterations, mitochondrial dysfunction, calcium dysregulation, oxidative stress, and therapeutic interventions contain significant overlap and repetition. Streamlining some background sections and reallocating space to a more detailed discussion of aging-specific mechanisms would help maintain focus and improve readability. In particular, the manuscript would benefit from expanding the sections on aging-induced MAM remodeling, age-dependent changes in MAM composition and ultrastructure, and the potential of MAM-targeted interventions as therapeutic strategies for sarcopenia.

      Finally, the review would gain from a stronger future perspectives section. Several important questions remain unresolved, including whether MAM disruption is a primary driver of muscle aging or a secondary consequence of mitochondrial dysfunction, how MAM architecture differs among muscle fiber types during aging, and whether MAM-associated proteins could serve as reliable biomarkers or therapeutic targets in human sarcopenia. Highlighting these knowledge gaps would further enhance the review's impact.

      Overall, the manuscript covers an important and emerging area of research and contains a valuable compilation of the relevant literature. Nevertheless, substantial revision is required to reduce the emphasis on well-established background information, deepen and critically analyze the aging-specific sections, and provide a more focused discussion of the role of MAMs in skeletal muscle aging and sarcopenia.

    1. eLife Assessment

      This study investigates the cellular mechanisms underlying theta-nested gamma oscillations in the medial entorhinal cortex; the experiments are rigorous, and the analyses and modeling provide potentially useful insights into cell-type-dependent circuit dynamics. However, the evidence supporting several key conclusions remains incomplete. The study is limited by conceptual constraints in the experimental design and a modeling approach that does not fully address underlying physiological mechanisms. Overall, this is a careful study that addresses how distinct neuronal populations in superficial MEC participate in theta-gamma coordination and provides new data linking cell-type-specific activity patterns to oscillatory network structure.

    2. Reviewer #1 (Public review):

      Summary:

      The question posed on cell-type-dependent relationships to theta-nested gamma rhythms is an important one. The authors use a variety of ontogenetic, imaging, electrophysiology, and computational techniques to show that reciprocal interactions between excitatory neurons and interneurons in the medial entorhinal cortex generate gamma oscillations. They measure LFP gamma, gamma power of postsynaptic currents in different neurons, spike phases with reference to LFP gamma, and spatial correlations of membrane potentials across a large population of neurons. Arguing (correctly) that gamma rhythm in this setting is generated through a pyramidal-interneuron network gamma (PING) mechanism, they demonstrate cell-type-specific differences in gamma phase-locking. While they show spatial dependencies of sub-threshold voltages and even argue for topographic clustering, these could simply be reflections of the synchronous stimulation paradigm that they use.

      Overall, I appreciate the methodology and rigor, but would have expected more from the study in terms of relevance to physiological stimulation conditions as well as in terms of mechanisms underlying the differences that they report here..

      Strengths:

      The authors are rigorous in how they conduct the experiments, report the data, and perform the analyses. The modeling respects the heterogeneities and is truthful to the experimental design. The conclusions on PING mechanisms are fine, but are not unexpected given the circuitry of the mEC.

      Weaknesses:

      The interpretation of the conclusions, while for the most part is fine, could have been better, especially given the conceptual limitations of the experimental design. The modeling part could have gone beyond simple descriptive matching and addressed mechanistic questions.

    3. Reviewer #2 (Public review):

      In this manuscript, the authors studied the cellular mechanism of theta-nested gamma oscillations in the medial entorhinal cortex (MEC) in vitro. The theta-nested gamma activity was induced by theta-modulated optogenetic stimulation of CaMKII+ neurons. In Figures 1 through 4, they describe the firing phase, synaptic input, and LFP-IPSC coupling of stellate cells, pyramidal cells, and interneurons. They then conducted voltage imaging, capturing the simultaneous activity of 41 cells, and found that subthreshold membrane potentials cluster in a weakly distance-dependent manner (Figure 5). The experiments and analysis are done rigorously for the most part.

      However, the results described in Figures 1 to 4 are largely descriptive and highly similar to those in their recent publication, which utilized almost identical experiments. While the voltage imaging data during theta-nested gamma oscillations are novel, the authors report data from only a single experiment, leaving it unclear whether the results are reproducible. Furthermore, without a comparison to in vivo data, it remains unclear what novel insights this manuscript provides to advance our understanding of the cellular mechanisms underlying theta-nested gamma oscillations.

      (1) The authors recently published another paper on the topic of theta-nested gamma oscillations in the MEC (Williams et al., eNeuro, 2026). In that study, they utilized a Thy1 promoter instead of the CaMKII promoter used here. The motivation for testing the CaMKII promoter in the current manuscript, as well as the novel insights expected from this experimental setup, remains unclear. Given that existing literature suggests inhibitory MEC cells play a critical role in theta activity (e.g., Gonzalez-Sulser et al., 2014)-implying that theta modulation should drive inhibitory rather than excitatory cells-the previous use of the Thy1 promoter appears closer to in vivo conditions than the CaMKII promoter used here.

      The overall conclusion of the current manuscript is that excitatory-inhibitory (E-I) interactions dominate the generation of theta-nested gamma oscillations. However, in their previous eNeuro paper, the authors demonstrated that the interneuron network gamma (ING) mechanism can sustain gamma oscillations without excitatory synaptic transmission. It seems expected that excitatory cells would be involved when the optogenetic stimulation selectively drives excitatory cells. If CaMKII stimulation is less physiological and artificially forces the theta-nested gamma activity to rely on excitatory connections, this conclusion could be misleading. It may potentially describe a mechanism that is irrelevant to physiological processes in vivo. Please see my comment 3, which is related to this point.

      In addition, Figures 1 and 2 heavily overlap with the authors' previous eNeuro publication. The differences in experimental settings and the motivation for performing almost identical experiments must be clearly articulated prior to these figures to avoid confusion. The authors must also justify why it is necessary to present such similar data, and explicitly point out the novel findings in the current paper compared to their previous work.

      (2) Using voltage imaging to investigate theta-nested gamma oscillations is novel. However, the impact of the findings from this experiment appears minimal in the manuscript's current state. The most novel and interesting observation is likely presented in Figure 6, where the authors identified clustered voltage correlations. However, this appears to be an n=1 experiment, and these findings should be replicated at least in a few experiments. Furthermore, the manuscript lacks a discussion or interpretation of this observation, making it unclear whether the result is biologically meaningful. Please find specific suggestions regarding this point below.

      (3) The authors' primary motivation for investigating the mechanisms underlying theta-modulated gamma oscillations is their potential role in grid cell firing. Therefore, it is critical that the mechanisms studied here in vitro accurately reflect in vivo processes. For this reason, greater effort should be made to better link this in vitro study with existing in vivo data. Numerous public in vivo datasets are available that detail the firing activity of putative principal cells and interneurons during exploratory behavior in mice. Intracellular recordings in awake animals have also been published, some of which the authors already cite. The data presented in Figures 1 and 4, for example, could be straightforwardly compared with those existing in vivo metrics. Furthermore, available in vivo silicon probe recordings could provide a reliable estimate of the spatial distribution of gamma-related spike activity. Such data should be compared with the voltage imaging results presented in this study.

      This limitation connects back to the first point. In this manuscript, the authors tested a different method for inducing theta-nested gamma oscillations (via the CaMKII promoter) than in their recent eNeuro paper (via the Thy1 promoter). The outcomes of these two induction methods must be systematically compared against in vivo data to determine which approach aligns more closely with physiological conditions. Without such a comparison, the scientific justification for testing a different promoter in this study remains unclear.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Williams et al. combine optogenetics, whole-cell electrophysiology, local field potential recordings, large-scale voltage imaging, and computational modeling to investigate the cellular and circuit mechanisms underlying theta-nested gamma oscillations in superficial medial entorhinal cortex (mEC). The authors propose that fast-spiking interneurons receive strong gamma-frequency excitatory drive and provide rhythmic inhibition onto principal neurons, supporting a pyramidal-interneuron network gamma (PING) mechanism. They further report cell-type-specific differences in gamma phase locking, spatial clustering of subthreshold voltage signals, and a network model reproducing several observed features, including interneuron bursting and gamma-cycle skipping in excitatory neurons.

      Strengths:

      The study is technically sophisticated and addresses an important question in entorhinal circuit function. The combination of intracellular recordings, voltage imaging, and computational modeling is a clear strength.

      Weaknesses:

      Several key conclusions developed from experimental results require additional raw data, statistical support, clearer methodological description, and more cautious interpretation. The computational modeling focuses primarily on stellate cells, whereas the experimental results suggest an important role for pyramidal neurons in PING dynamics. This creates inconsistency between theory and experiments.

    5. Author response:

      We thank the editors and reviewers for their thoughtful comments. Below, we list our provisional responses to the reviewers’ major points:

      On the rationale for CaMKIIα versus Thy1-driven stimulation and physiological relevance: We agree that we did not make clear the motivation for using CaMKIIα-driven stimulation, distinct from the Thy1-driven paradigm in our previous work (Williams et al., 2026). Using the Thy1 driver, both excitatory and inhibitory cells received direct theta drive. In contrast, CaMKIIα expression is largely restricted to principal neurons. Comparing these models lets us isolate a "driven I-cell" PING mechanism from the "E cell recovers first" mechanism relevant when interneurons are also directly driven.

      Regarding physiological relevance, Gonzalez-Sulser et al. (2014) found that septal GABAergic projections selectively and directly inhibit mEC interneurons, rather than exciting either principal cells or interneurons, implying that theta drive in vivo likely acts through rhythmic disinhibition of interneurons rather than direct excitation of any cell type. Neither the Thy1 nor the CaMKIIα paradigm reproduces this disinhibitory mechanism: both rely on excitatory optogenetic drive rather than rhythmic inhibition of interneurons, and replicating the natural drive (tonic excitatory tone plus rhythmic, interneuron-selective inhibition) is technically difficult in acute slices, which are largely quiescent without exogenous stimulation. We therefore view CaMKIIα and Thy1 as complementary approximations, each isolating a different circuit interaction. If forced to choose, we’d argue that the CaMKIIα is a better model of disinhibition of excitatory neurons. We will revise the Discussion regarding this point.

      On reproducibility of the voltage imaging findings: We thank the reviewer for this comment and agree that clarification is warranted.

      The voltage imaging dataset combines two levels of analysis with different sample sizes. The population-level firing and spike-correlation analyses (Fig. 5F–H) are pooled across multiple imaging sessions (n = 240 neurons). The spatial clustering analysis of subthreshold voltage correlations (Fig. 6, and the corresponding example traces in Fig. 5A–E) are drawn from a single representative recording session, as the reviewer correctly notes. We have voltage imaging data from 14 fields of view (1 FOV per slice) across 6 mice (240 neurons total; 3–41 neurons per FOV). In revision, we will extend the clustering and spatial-correlation analysis from Fig. 6 across sessions to assess whether the reported organization is reproducible, rather than relying on a single example. We will also revise the text to distinguish clearly which analyses are single-session versus pooled.

      On restricting the computational model of excitatory neurons to stellate cells: We modeled stellate cells as the excitatory population because they are the principal cells reciprocally connected to fast-spiking PV+ interneurons (Fuchs et al., 2016), the interneuron class most directly implicated in theta-nested gamma. Pyramidal cells, by contrast, are primarily connected via 5-HT3a-positive interneurons (Fuchs et al., 2016), with the exception of a subset of "intermediate" pyramidal cells that do show reciprocal PV+ connectivity. Our model, which captures the full measured heterogeneity of stellate cell and PV+ interneuron intrinsic properties and their reciprocal connectivity, is, to our knowledge, the most biophysically constrained implementation of this specific microcircuit to date. Incorporating the PV+-connected intermediate pyramidal population is a natural next step. Because this refinement, which requires more experimental data, is nontrivial and beyond the scope of this study, we will note this explicitly as a limitation of the current model in the revised Discussion.

      In vivo comparison (temporal/phase-locking): We agree that grounding our findings in existing in vivo data strengthens the study and will add these comparisons to the revision.

      Our whole-cell recordings reproduce the temporal organization in vivo and provide further insights into cell-type differences between the principal cells. All cell types were strongly phase-locked to theta, while gamma phase-locking declined across successive spikes, with stellate cells decoupling after the first spike and pyramidal cells after the second. This earlier decoupling in stellate cells may contribute to their weaker theta rhythmicity reported in freely moving rats (Ray et al., 2014; Tang et al., 2014). In extracellular recordings from behaving mice, spike-train cross-correlation identifies putative monosynaptic excitatory connections (1–4 ms) from principal cells onto fast-spiking interneurons (Latuske et al., 2015); the excitation-to-inhibition offset we measured is of comparable magnitude, here resolved as a synaptic-current delay in electrophysiologically classified cell types.

      We note that bursting and theta engagement have been assigned inconsistently across in vivo datasets. Bursty cells are preferentially classified as putative stellate by spikepattern classifiers (Latuske et al., 2015), while anatomically identified pyramidal cells are reported as the bursty, theta-rhythmic population in other work (Ebbesen et al., 2016). Because our cell-type assignments are based on subthreshold intrinsic properties (membrane sag, time constant) rather than spike patterning, our phase-locking results are independent of this classification ambiguity.

      In vivo comparison (spatial organization): We agree high-density silicon-probe datasets are the appropriate reference here. To our knowledge, the anatomical distribution of gamma-locked spiking in superficial mEC has not been characterized in vivo. The highest-density available recordings (Gardner et al., 2022) analyze population activity in the decoded state rather than tissue coordinates, do not examine gamma, and are restricted to grid cells. We regard the dissociation we observe between spatially clustered subthreshold input and spatially distributed spiking as a principal advance of the present study, and as a testable prediction for future high-density recordings.

      Ebbesen CL, Reifenstein ET, Tang Q, Burgalossi A, Ray S, Schreiber S, Kempter R, Brecht M. 2016. Cell Type-Specific Differences in Spike Timing and Spike Shape in the Rat Parasubiculum and Superficial Medial Entorhinal Cortex. Cell Reports 16:1005–1015. DOI: https://doi.org/10.1016/j.celrep.2016.06.057

      Fuchs EC, Neitz A, Pinna R, Melzer S, Caputi A, Monyer H. 2016. Local and Distant Input Controlling Excitation in Layer II of the Medial Entorhinal Cortex. Neuron 89:194–208. DOI: https://doi.org/10.1016/j.neuron.2015.11.029

      Gardner RJ, Hermansen E, Pachitariu M, Burak Y, Baas NA, Dunn BA, Moser M-B, Moser EI. 2022. Toroidal topology of population activity in grid cells. Nature 602:123–128. DOI: https://doi.org/10.1038/s41586-021-04268-7

      Gonzalez-Sulser A, Parthier D, Candela A, McClure C, Pastoll H, Garden D, Sürmeli G, Nolan MF. 2014. Gabaergic projections from the medial septum selectively inhibit interneurons in the medial entorhinal cortex. Journal of Neuroscience 34:16739–16743. DOI: https://doi.org/10.1523/JNEUROSCI.1612-14.2014, PMID: 25505326

      Latuske P, Toader O, Allen K. 2015. Interspike Intervals Reveal Functionally Distinct Cell Populations in the Medial Entorhinal Cortex. Journal of Neuroscience 35:10963–10976. DOI: https://doi.org/10.1523/JNEUROSCI.0276-15.2015

      Ray S, Naumann R, Burgalossi A, Tang Q, Schmidt H, Brecht M. 2014. Grid-Layout and Theta-Modulation of Layer 2 Pyramidal Neurons in Medial Entorhinal Cortex. Science 343:891–896. DOI: https://doi.org/10.1126/science.1243028

      Tang Q, Burgalossi A, Ebbesen CL, Ray S, Naumann R, Schmidt H, Spicher D, Brecht M. 2014. Pyramidal and Stellate Cell Specificity of Grid and Border Representations in Layer 2 of Medial Entorhinal Cortex. Neuron 84:1191–1197. DOI: https://doi.org/10.1016/j.neuron.2014.11.009

      Williams B, Vedururu Srinivas A, Baravalle R, Fernandez FR, Canavier CC, White JohnA. 2026. Fast spiking interneurons autonomously generate fast gamma oscillations in the medial entorhinal cortex with excitation strength tuning ING–PING transitions. eneuro ENEURO.0452-25.2026. DOI: https://doi.org/10.1523/ENEURO.0452-25.2026

    1. eLife Assessment

      This study makes a solid and valuable contribution to elucidating the intricate relationship between mitochondrial calcium and neuronal survival. Well-controlled experiments show that homeostatic mitochondrial calcium correlates with the most resilient neuronal subtypes after optic nerve injury. However, altering mitochondrial calcium levels does not affect neuronal survival as initially predicted by this correlation.

    2. Reviewer #1 (Public review):

      Summary:

      This study investigates how Ca2+ levels inside the RGCs' mitochondria relate to whether these cells survive or die after injury to the optic nerve. The authors used advanced in vivo fundus live imaging techniques in mice to watch these changes unfold in real time, combined with genetic and drug-based tools to alter calcium flow into these compartments. Their central finding is a striking paradox: cells that naturally survive injury tend to have higher baseline calcium levels in these compartments, yet experimentally reducing calcium entry protects the broader population of cells from death.

      Strengths:

      The authors are applying sophisticated biosensors to track cellular chemistry in living animals over days and weeks. The tools and methods are creative and direct to detect the longitudinal RGC degeneration with mito-Ca2+ imaging. The topic and research aspect are novel and attractive. The results are significant, showing a clear relationship between the mito-Ca2+ regulatory machinery and cell survival.

      Weaknesses:

      The details of the mitochondrial-located signal of the Ca2+ sensor need to be further proved in the mito-matrix or between the mito-membranes. The study primarily describes a correlation and a surprising experimental outcome without fully explaining the underlying biological reasons for the paradox. While the evidence supporting the phenomenon is good, the mechanistic insight into why high calcium is linked to survival, or why lowering it helps after injury, remains limited.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript by McCraken and colleagues provides a continuation of their 2023 study (Cell Reports 42:113165) characterizing calcium regulation in retinal ganglion cells (RGCs) after acute optic nerve damage (a 10s crush using an intraorbital approach). This work is principally focused on how mitochondrial calcium stores change in both RGCs that are resilient and susceptible to injury. They report that resilient RGCs typically exhibited high calcium levels, but paradoxically, manipulating mitoCa2+ levels was more protective when the stores were reduced. Overall, regardless of susceptibility, mitoCa2+ levels decreased after injury, which is opposite to other reports that mitoCa2+ increases in degenerating neurons. The manipulation of mitoCa2+ was conducted both pharmacologically (Ru265) and by overexpression or knockdown of a primary calcium uniporter MCU. The evaluation of mitoCa2+ was conducted by using a reporter (Twitch2b) that was targeted to the mitochondria.

      Strengths:

      Many of the experiments are elegant and well-performed.

      Weaknesses:

      (1) Some experiments require further controls to validate that reagents are doing what they are intended to do.

      (2) Some findings can have alternate interpretations that are not considered.

      (3) There is a broad generalization to the biology of all RGCs that may not be biologically relevant to different RGC subtypes.

    4. Reviewer #3 (Public review):

      Summary:

      Following previous work that demonstrated a relationship between higher homeostatic cytosolic calcium and lower retinal ganglion cell (RGC) apoptosis following injury to their axons, McCracken et al. investigated whether homeostatic calcium levels of the endoplasmic reticulum (ER) or mitochondria provide additional insights into the mechanisms by which calcium influences RGC survival. Their study reveals that homeostatic mitochondrial calcium shows a similar positive correlation with RGC survival. Despite that correlation, pharmacologic or genetic methods to lower mitochondrial calcium improved, rather than reduced, the survival of injured RGCs, while a genetic approach intended to increase mitochondrial calcium resulted in more RGC loss. These findings highlight the complexities of calcium regulation in modulating neuronal survival and raise important questions of how homeostatic levels of mitochondrial calcium affect stress responses that themselves can be either neuroprotective or neurodegenerative.

      Strengths:

      This study tackles an intriguing hypothesis that differences in calcium ion homeostasis in specific organelles may contribute to differences in survival of various RGC subtypes after optic nerve injury. This is a technically demanding question, and a primary strength of this work is its attention to, and meticulous reporting of, appropriate controls and, where applicable, seemingly contradictory results. Among these are careful evaluation of the effects of drug (or vehicle) delivery and genetic manipulations with and without injury and over extended time courses. The combination of thoughtful pharmacologic and genetic approaches makes for a thorough analysis of a challenging set of questions. The result is a study that provides a helpful perspective on the complicated roles that calcium, and especially mitochondrial calcium, can play across neuronal insults, neuronal types, and neuronal subtypes.

      Weaknesses:

      Given the paradoxical results, it would be helpful to have a clearer picture of how strongly the overexpression and knockdown of MCU altered the mitochondrial calcium levels. There may be potential for extraordinarily strong effects that would need to be tuned by using different shRNAs or promoters to more closely align with the observed differences between surviving RGCs and those that die. The investigation includes a relatively small number of resilient RGC subtypes, using the markers SPP1 and TBR2, raising questions of how generalizable the trend is between mitochondrial calcium levels and RGC resilience. The analysis and implications of Figure 3D might benefit from including not only the provided 50:50 split between "high" and "low" but also views of the data after splitting into thirds, fourths, and perhaps even fifths. The authors' inference that higher homeostatic calcium in more resilient RGCs may result in chronic mitochondrial stress is intriguing and worthy of more experimental investigation than is currently provided.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study investigates how Ca2+ levels inside the RGCs' mitochondria relate to whether these cells survive or die after injury to the optic nerve. The authors used advanced in vivo fundus live imaging techniques in mice to watch these changes unfold in real time, combined with genetic and drug-based tools to alter calcium flow into these compartments. Their central finding is a striking paradox: cells that naturally survive injury tend to have higher baseline calcium levels in these compartments, yet experimentally reducing calcium entry protects the broader population of cells from death.

      Strengths:

      The authors are applying sophisticated biosensors to track cellular chemistry in living animals over days and weeks. The tools and methods are creative and direct to detect the longitudinal RGC degeneration with mito-Ca2+ imaging. The topic and research aspect are novel and attractive. The results are significant, showing a clear relationship between the mito-Ca2+ regulatory machinery and cell survival.

      Weaknesses:

      The details of the mitochondrial-located signal of the Ca2+ sensor need to be further proved in the mito-matrix or between the mito-membranes. The study primarily describes a correlation and a surprising experimental outcome without fully explaining the underlying biological reasons for the paradox. While the evidence supporting the phenomenon is good, the mechanistic insight into why high calcium is linked to survival, or why lowering it helps after injury, remains limited.

      We appreciate Reviewer #1’s assessment of our manuscript. We also agree that we should have more clearly indicated that our mitochondrial Ca2+ sensor (Cox8-Twitch2b) is localized to the mitochondrial matrix. The Cox8-mitochondrial localization peptide is a well-established tool first identified in 1992 by Rizzuto and colleagues (Rizzuto, Simpson and Pozzan, 1992). We should have cited this work in our manuscript and will add it to our references. Further, as discussed in our submission, Cox8-Twitch2b has previously been validated for mitochondrial Ca2+ measurements in CNS axons (Witte et al., 2019). Thus, given the decades of use and characterization for this toolset, and the fact that we have pharmacological data supporting mitochondrial matrix localization of Cox8-Twitch2b, we do not feel it is strongly necessary to demonstrate mitochondrial matrix versus inner membrane space localization. However, we could attempt immuno-electron microscopy if this is deemed critical.

      We also agree that the mechanism by which reducing mitochondrial Ca2+ is protective would be satisfying and strengthen this study. But we feel it is beyond the scope of this project. It is likely manifold since mitochondrial Ca2+ impacts many vital cellular functions relevant to pathology including metabolism and apoptosis. We ultimately believe that an adequate investigation of these mechanisms would significantly slow down the dissemination of the core novel findings presented herein.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by McCraken and colleagues provides a continuation of their 2023 study (Cell Reports 42:113165) characterizing calcium regulation in retinal ganglion cells (RGCs) after acute optic nerve damage (a 10s crush using an intraorbital approach). This work is principally focused on how mitochondrial calcium stores change in both RGCs that are resilient and susceptible to injury. They report that resilient RGCs typically exhibited high calcium levels, but paradoxically, manipulating mitoCa2+ levels was more protective when the stores were reduced. Overall, regardless of susceptibility, mitoCa2+ levels decreased after injury, which is opposite to other reports that mitoCa2+ increases in degenerating neurons. The manipulation of mitoCa2+ was conducted both pharmacologically (Ru265) and by overexpression or knockdown of a primary calcium uniporter MCU. The evaluation of mitoCa2+ was conducted by using a reporter (Twitch2b) that was targeted to the mitochondria.

      Strengths:

      Many of the experiments are elegant and well-performed.

      Weaknesses:

      (1) Some experiments require further controls to validate that reagents are doing what they are intended to do.

      We agree with Reviewer #2 that our AAV manipulations of shMCU and MCU overexpression should be analyzed to verify how they alter mitochondrial Ca2+. To do this, we will co-express gene therapy vectors to lower and raise MCU expression with mito-Twitch2b biosensor and perform direct measurements of mitochondrial Ca2+. We will then determine if there is a relationship between gene expression level (inferred by mCherry intensity) and mitochondrial Ca2+ within samples, and if mean mitochondrial Ca2+ levels in treatments are higher or lower than mCherry reporter only controls.

      (2) Some findings can have alternate interpretations that are not considered.

      We will expand our Results and Discussion sections to broaden the interpretations of our data.

      (3) There is a broad generalization to the biology of all RGCs that may not be biologically relevant to different RGC subtypes.

      We agree that a more fine-grained understanding of RGC mitochondrial Ca2+ diversity would make interpretations of our data stronger. In our revisions, we will thus expand the number of RGC families in which we directly measure homeostatic mitochondrial Ca2+ levels. To do this, we will perform in vivo mito-Twitch2b measurements, collect and fix retinal wholemounts and immunostain for ON-OFF-direction selective RGCs using the marker CART and F-RGCs using the marker Foxp2. This will provide a complement of well-surviving RGC types (alpha and intrinsically photosensitive RGCs already examined) and poorly-surviving types.

      Reviewer #3 (Public review):

      Summary:

      Following previous work that demonstrated a relationship between higher homeostatic cytosolic calcium and lower retinal ganglion cell (RGC) apoptosis following injury to their axons, McCracken et al. investigated whether homeostatic calcium levels of the endoplasmic reticulum (ER) or mitochondria provide additional insights into the mechanisms by which calcium influences RGC survival. Their study reveals that homeostatic mitochondrial calcium shows a similar positive correlation with RGC survival. Despite that correlation, pharmacologic or genetic methods to lower mitochondrial calcium improved, rather than reduced, the survival of injured RGCs, while a genetic approach intended to increase mitochondrial calcium resulted in more RGC loss. These findings highlight the complexities of calcium regulation in modulating neuronal survival and raise important questions of how homeostatic levels of mitochondrial calcium affect stress responses that themselves can be either neuroprotective or neurodegenerative.

      Strengths:

      This study tackles an intriguing hypothesis that differences in calcium ion homeostasis in specific organelles may contribute to differences in survival of various RGC subtypes after optic nerve injury. This is a technically demanding question, and a primary strength of this work is its attention to, and meticulous reporting of, appropriate controls and, where applicable, seemingly contradictory results. Among these are careful evaluation of the effects of drug (or vehicle) delivery and genetic manipulations with and without injury and over extended time courses. The combination of thoughtful pharmacologic and genetic approaches makes for a thorough analysis of a challenging set of questions. The result is a study that provides a helpful perspective on the complicated roles that calcium, and especially mitochondrial calcium, can play across neuronal insults, neuronal types, and neuronal subtypes.

      Weaknesses:

      Given the paradoxical results, it would be helpful to have a clearer picture of how strongly the overexpression and knockdown of MCU altered the mitochondrial calcium levels. There may be potential for extraordinarily strong effects that would need to be tuned by using different shRNAs or promoters to more closely align with the observed differences between surviving RGCs and those that die. The investigation includes a relatively small number of resilient RGC subtypes, using the markers SPP1 and TBR2, raising questions of how generalizable the trend is between mitochondrial calcium levels and RGC resilience. The analysis and implications of Figure 3D might benefit from including not only the provided 50:50 split between "high" and "low" but also views of the data after splitting into thirds, fourths, and perhaps even fifths. The authors' inference that higher homeostatic calcium in more resilient RGCs may result in chronic mitochondrial stress is intriguing and worthy of more experimental investigation than is currently provided.

      We agree with the feedback from Reviewer #3, especially as it aligns with input from other reviewers. As these points agree with aspects above we will briefly reiterate our proposed revisions. We will validate the true effects on mitochondrial Ca2+ levels after gene therapy treatments by co-injecting AAV-mito-Twitch2b and AAV-shMCU or AAV-MCU. We will measure mitochondrial Ca2+ levels and correlate these levels with mCherry reporter expression intensity to determine the effect size of these treatments, and compare sample mean mitochondrial Ca2+ levels with those of mCherry control AAV.

      To further map the variance in homeostatic mitochondrial Ca2+ levels to RGC types we will perform in vivo mito-Twitch2b imaging, and then immunostain for ON-OFF-direction selective RGCs (CART) and F-RGCs (Foxp2), two poorly surviving RGC types.

      Lastly, we agree with Reviewer #3 that finer delineation between mitochondrial Ca2+ levels and their relationship to survival may be informative. We will split RGCs into smaller subgroups based on homeostatic mitochondrial Ca2+ levels and examine their survival outcome.

      Overall, we thank the Reviewers for their feedback, and believe the suggested changes will greatly strengthen our study.

      REFERENCES

      Rizzuto R., Simpson A.W. and Pozzan T. (1992). Rapid changes of mitochondrial Ca2+ revealed by specifically targeted recombinant aequorin. Nature, 358 (6384): 325-327.

      Witte M.E., Schumacher A-M., Mahler C.F., Bewersdorf J.P., Lehmitz J., Scheiter A., Sanchez P., Williams P.R., Griesbeck O., Naumann R., Misgeld T. and Kerschensteiner M. (2019). Calcium influx through plasma-membrane nanoruptures drives axon degeneration in a model of multiple sclerosis. Neuron, 101(4): 615-624.

    1. eLife Assessment

      This article describes the comprehensive metabolic phenotype of a mouse model of Down Syndrome, together with supporting transcriptomic, metabolomic, and biochemical data. The evidence presented is compelling and highlights several core phenotypes including insulin resistance, dyslipidemia, and tissue signatures indicating inflammatory and cellular stress pathways. Similarities and differences in male and female mice are highlighted. This important study provides essential groundwork for the further genetic dissection of dosage-sensitive genes causing metabolic dysregulation in Down Syndrome.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      Chen et al. describe metabolic phenotypes in Dp16 Down Syndrome mice, specifically the Dp(16)1Yey/+ mice - segmental duplication model carrying a majority of the triplicated Hsa21 gene orthologs. The group has performed metabolic phenotyping data in chow and high-fat diets, as well as undertaking a transcriptomic and metabolomic approach in tissues such as white and brown adipose tissues, liver, skeletal muscle, and hypothalamus to reveal both shared and sex-specific differences. The group describes sexual dimorphism in body weight, body temperature, food intake, and physical activity. Core shared features are insulin resistance, glucose intolerance, impaired lipid clearance, and dyslipidaemia in the Dp16 mice. They report tissue signatures of immune activation and a pro-inflammatory state, ER and oxidative stress, fibrosis, impaired glucose and fatty acid catabolism, altered lipid and bile acid profiles, and reduced mitochondrial respiration in Dp16 mice.

      Strengths:

      Overall, this is a good study with detailed, comprehensive data from an excellent group who have previously published on metabolic phenotyping of 2 other Down Syndrome mouse models. Although somewhat descriptive, it does certainly add to the current field and understanding of strengths and weaknesses of Down Syndrome mouse models, as well as identifying new features whilst strengthening previously suggested mechanisms.

    3. Reviewer #2 (Public review):

      Summary:

      Human DS is associated with metabolic dysfunction in humans, but the precise details of this have not been studied in detail. Here, the authors use a mouse model of DS to study systemic metabolic and transcriptional responses in key metabolic tissues to provide a deep understanding of the metabolic changes associated with DS. As part of his work, the authors also aimed to help inform the selection of a mouse model that best reflects the metabolic profile of DS, through comparison with other DS model metabolic data.

      The data presented in this model will be of interest to those in the field of metabolism. The immediate impact is unclear, but the breadth of data presented makes this a very useful resource.

      Strengths:

      (1) This work builds on other comprehensive analyses that the authors have performed in other DS mouse models.

      (2) The authors note common metabolic disturbances between male and female mice (e.g., insulin resistance) alongside clearly sexually dimorphic phenotypes (e.g., body weight). Studying both sexes in this context is important.

      (3) The authors have written the paper in a way that integrates a large number of observations well. There is complex data, and a high degree of sexual dimorphism. The study has generated a valuable and wide-ranging dataset comprising molecular, biochemical, and physiological data that will be useful for further, more mechanistic studies of metabolism in DS.

      (4) For specific observations, like the findings of altered body temperature in male and female mice, the authors undertake follow-up hypothesis-driven analyses of BAT mitochondria and specific hormones. Although these analyses do not explain the change in temperature, they ensure the study is not purely descriptive in nature.

    4. Reviewer #3 (Public review):

      Summary:

      The article by Chen et al. describes the comprehensive metabolic profiling of DP16 mice, a Down syndrome model that carries a duplicated segment of the mouse chromosome syntenic to human chromosome 21. The authors note that this model is superior to previously used models, based on genetics, as ~65% of the chromosome 21 orthologues. The metabolic phenotypes also appear to be more consistent with those observed in humans with Down Syndrome. The study lays the groundwork for a more detailed genetic dissection of dosage-sensitive genes that contribute to the metabolic deficits observed in Down Syndrome.

      Strengths:

      There is an enormous amount of data in this manuscript, and the methods are described with adequate attention to detail. A strength of the manuscript is that both male and female mice were analyzed, so that concordant and discordant phenotypes were identified. Both males and females had evidence of insulin resistance. Transcriptomic and metabolomic data revealed impaired pathways for lipid metabolism, a pro-inflammatory state, reduced mitochondrial health and oxidative stress. Although the effects of a high-fat diet on weight gain were divergent, this diet caused worsened insulin resistance in both males and females.

      The discussion is excellent. Limitations of the study are well described. This reviewer does not identify any critical missing data.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Chen et al. describe metabolic phenotypes in Dp16 Down Syndrome mice, specifically the Dp(16)1Yey/+ mice - segmental duplication model carrying a majority of the triplicated Hsa21 gene orthologs. The group has performed metabolic phenotyping data in chow and high-fat diets, as well as undertaking a transcriptomic and metabolomic approach in tissues such as white and brown adipose tissues, liver, skeletal muscle, and hypothalamus to reveal both shared and sex-specific differences. The group describes sexual dimorphism in body weight, body temperature, food intake, and physical activity. Core shared features are insulin resistance, glucose intolerance, impaired lipid clearance, and dyslipidaemia in the Dp16 mice. They report tissue signatures of immune activation and a pro-inflammatory state, ER and oxidative stress, fibrosis, impaired glucose and fatty acid catabolism, altered lipid and bile acid profiles, and reduced mitochondrial respiration in Dp16 mice.

      Strengths:

      Overall, this is a good study with detailed, comprehensive data from an excellent group who have previously published on metabolic phenotyping of 2 other Down Syndrome mouse models. Although somewhat descriptive, it does certainly add to the current field and understanding of strengths and weaknesses of Down Syndrome mouse models, as well as identifying new features whilst strengthening previously suggested mechanisms.

      Weaknesses:

      Many aspects of this study have been described in other Down syndrome mouse models, though there are certainly aspects that are new. It would be useful if the authors could do a direct critique and comparison with previous publications in the area, utilizing the same Down Syndrome mouse model. There are also a few limitations in the number of animals used and the interpretation of the data that should be acknowledged.

      We have cited all relevant publications using Down syndrome mouse models. Regarding the Dp16 model, we have cited and discussed the only other study addressing metabolic aspects beyond body weight (Reference #138; PMID: 39803786). While that study reported glucose intolerance, insulin resistance, and defective insulin secretion, we did not measure pancreatic insulin content in our mice. Crucially, while the previous study found no sexual dimorphism, our study observed extensive sexual dimorphism in body weight gain, tissue-specific gene expression, and serum and liver metabolite changes.

      Regarding sample size, we used 6 mice per genotype per sex for transcriptomic and metabolomic analyses; this is constrained by the cost of performing these omics-type analyses. For mitochondrial respiration assays, we used 9–10 mice, and for most other in vivo and ex vivo assays, we utilized 12–15 mice, with some assays exceeding 20. We believe these sample sizes are robust and appropriate for this study.

      Reviewer #2 (Public review):

      Summary:

      Human DS is associated with metabolic dysfunction in humans, but the precise details of this have not been studied in detail. Here, the authors use a mouse model of DS to study systemic metabolic and transcriptional responses in key metabolic tissues to provide a deep understanding of the metabolic changes associated with DS. As part of his work, the authors also aimed to help inform the selection of a mouse model that best reflects the metabolic profile of DS, through comparison with other DS model metabolic data.

      The data presented in this model will be of interest to those in the field of metabolism. The immediate impact is unclear, but the breadth of data presented makes this a very useful resource.

      Strengths:

      (1) This work builds on other comprehensive analyses that the authors have performed in other DS mouse models.

      (2) The authors note common metabolic disturbances between male and female mice (e.g., insulin resistance) alongside clearly sexually dimorphic phenotypes (e.g., body weight). Studying both sexes in this context is important.

      (3) The authors have written the paper in a way that integrates a large number of observations well. There is complex data, and a high degree of sexual dimorphism. The study has generated a valuable and wide-ranging dataset comprising molecular, biochemical, and physiological data that will be useful for further, more mechanistic studies of metabolism in DS.

      (4) For specific observations, like the findings of altered body temperature in male and female mice, the authors undertake follow-up hypothesis-driven analyses of BAT mitochondria and specific hormones. Although these analyses do not explain the change in temperature, they ensure the study is not purely descriptive in nature.

      Weaknesses:

      (1) Assessing metabolism using dynamic testing is a strength. ITT, GTT and LTTs are included.

      (2) The dosing for GTTs, ITTs and LTTs was performed per body weight. But the mice under chow and HFD had different body weights. This may compromise the interpretation of the data. Further, ITTs are presented as percentage change, and this can be heavily influenced by baseline glucose measures. The changes appear quite dramatic, so can the authors plot the raw data instead?

      We have updated the ITT data plots to show raw glucose values instead of percentage change. Regarding the dosing, we believe basing it on body weight is an appropriate approach. This method is consistent with nearly all published rodent studies, as blood volume and metabolic tissues such as skeletal muscle and adipose tissue scale with body weight. Adjusting for weight prevents potentially erroneous conclusions. As for the diet groups, we compared WT and Dp16 mice only within the same diet group (Chow or HFD) rather than across different diets. We believe this ensures a valid and appropriate comparison for our study.

      (3) In addition, throughout the manuscript, it is not clear which tissues are the most dominant in disrupting metabolism. The ITT and GTT are composite measures across tissues. Tissue-specific analyses using a clamp technique or isolated tissues may provide more clarity here.

      Our data suggest a systemic metabolic deficit across multiple tissues, supported by tolerance tests, pan-tissue transcriptomic analyses, and liver and serum metabolite profiling. This is consistent with the triplication of genes in Down syndrome, several of which have known metabolic roles as highlighted in our discussion. We do not have evidence to support the role of a dominant tissue that contributes to the systemic metabolic dysfunction.

      Regarding the suggestion to use a clamp technique, we agree this would effectively determine whether insulin resistance is localized in the liver or skeletal muscle. However, we do not currently have the necessary equipment at Johns Hopkins University to perform these experiments. Conducting this work would require sending separate cohorts of WT and Dp16 male and female mice (on both chow and HFD) to an NIH-funded Mouse Metabolic Phenotyping Centre (MMPC). While we appreciate the value of this approach, we believe such labor-intensive experimentation falls beyond the scope of the present study.

      (4) One of the aims of the study was "to help inform the selection of mouse model that best reflects the metabolic profile of DS". The discussion does not contain a comparison between the previous work on different strains and relative to known human data.

      We chose not to include a comparison of different mouse models in the "Discussion" section because we previously highlighted the widely used Down syndrome models (Ts65Dn, Tc1, and TcMAC21) and their associated caveats in the "Introduction." Given the significant limitations of those models such as hypermetabolism in TcMAC21 and the presence of 41 triplicated protein-coding genes unrelated to human chromosome 21 we focused our in-depth metabolic analyses on the Dp16 model, which does not share these issues. We felt that restating this information in the "Discussion" would be unnecessarily repetitive.

      (5) Data availability. Raw metabolomic data should be made available.

      We have uploaded all metabolomics data, along with details regarding sample processing and data analysis, to the Metabolomics Workbench, an NIH-funded public repository. We have updated the "Methods" and "Data Availability" sections of the manuscript to include this information and the corresponding access link.

      Reviewer #3 (Public review):

      Summary:

      The article by Chen et al. describes the comprehensive metabolic profiling of DP16 mice, a Down syndrome model that carries a duplicated segment of the mouse chromosome syntenic to human chromosome 21. The authors note that this model is superior to previously used models, based on genetics, as ~65% of the chromosome 21 orthologues. The metabolic phenotypes also appear to be more consistent with those observed in humans with Down Syndrome. The study lays the groundwork for a more detailed genetic dissection of dosage-sensitive genes that contribute to the metabolic deficits observed in Down Syndrome.

      Strengths:

      There is an enormous amount of data in this manuscript, and the methods are described with adequate attention to detail. A strength of the manuscript is that both male and female mice were analyzed, so that concordant and discordant phenotypes were identified. Both males and females had evidence of insulin resistance. Transcriptomic and metabolomic data revealed impaired pathways for lipid metabolism, a pro-inflammatory state, reduced mitochondrial health and oxidative stress. Although the effects of a high-fat diet on weight gain were divergent, this diet caused worsened insulin resistance in both males and females.

      The discussion is excellent. Limitations of the study are well described. This reviewer does not identify any critical missing data.

      Weaknesses:

      It might have been helpful to have included blood pressure measurements, given the differences in 19-Nor-deoxycorticosterone. The discussion references several articles that describe sex-dependent differences in metabolic phenotypes in humans with Down syndrome, and it might have been helpful to state more explicitly whether these differences correlate with those observed here in mice.

      We appreciate the suggestion of blood pressure measurements. While we agree this is an important metric, given the metabolic focus of the present study and the significant volume of data already presented, we feel that blood pressure analysis is beyond the current scope and better suited for a follow-up study.

      Our study highlights sex differences in metabolic phenotypes in individuals with Down syndrome. While most published human studies focus on a limited set of parameters such as body weight, adiposity, serum lipoprotein profile, and fasting lipid/glucose levels our mouse data remain generally concordant with these findings. Beyond these standard measurements, we also observed substantial sex differences in pan-tissue transcriptomes as well as serum and liver metabolites.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) A major question is how these findings compare to data that have previously been published. For example, Lamantia et al. Bone 2024 and Dard et al. European Journal of Pharmacology 2025 both report no changes in body weight using the same Dp(16)1Yey Down syndrome mouse model? There is also a recent publication on liver dysfunction in Down Syndrome using the same mouse model. It would be useful to understand some of the similarities and differences of what is being reported by Dunn et al. Cell Rep 2026. In this assessment, there is an in-serum alanine transaminase (ALT) level, which was not the case in Dunn et al?

      For the Lamantia et al. Bone 2024 study, the authors only measured the body weights of Dp16 mice at 6 weeks of age. Our findings at 6 weeks align with Lamantia et al., showing no weight differences between Dp16 and WT mice of either sex (Fig. 2A and C). For the Dard et al. 2025 study, the authors only measured the body weights of Dp16 mice at 12 weeks old (P90) and observed no differences in body weights between genotype of either sex. At 12 weeks of age, we also did not observe body weight differences between Dp16 male mice and WT littermates (Fig. 2A). However, at 12 weeks of age, the Dp16 female mice clearly gained more weight compared to WT littermates (Fig. 2C). Our study tracked weights weekly from 6 to 16 weeks, revealing that while Dp16 females start at weights similar to WT littermates, the groups diverge over time. The reason for the difference between our findings and the single-point measurement by Dard et al. is unclear. Notable variables include:

      Mouse Sourcing: We obtained all cohorts and littermate controls from Jackson Laboratory, while Dard et al. bred their mice in-house.

      Diet: We used Envigo standard chow (catalogue # 2018SX). Dard et al. did not specify the chow used in their study.

      It remains uncertain whether these or other environmental factors contribute to the observed weight differences in female mice.

      In the Dunn et al study (Cell Rep 2026), they also performed metabolic analyses on serum and liver tissue in Dp16 mice. Consistent with their metabolic analyses of serum and liver tissue in Dp16 mice, we also observed the upregulation of multiple bile acids, including taurochenodeoxycholic, tauromuricholic, taurolithocholic, and lithocholic acids. Furthermore, our findings align with theirs regarding the transcriptomic and biochemical signatures of hepatic inflammation and fibrosis. However, there are two notable differences between our studies:

      (1) Liver Injury Markers: We observed an elevation in serum ALT, whereas the Dunn et al. study did not.

      (2) Sex Differences: We identified significant sex differences in the Dp16 transcriptome and metabolome. In contrast, Dunn et al. reported minimal to no sex differences and consequently combined male and female data for all analyses.

      Because Dunn et al. combined male and female data, a sex-stratified comparison between our results (separated by sex) and theirs was not feasible.

      (2) It would be important to understand trends in wild-type animals compared to Dp16 mice. For example, the sex specific and non-specific features - are any of these described in obesogenic wild-type animals fed on a high-fat diet? I.e., are the same features at play and just exacerbated in Dp16, or is this a Dp16-specific feature of systemic metabolism?

      Published literature indicates that WT females typically gain significantly less weight on a high-fat diet (HFD) than WT males. However, our data suggest that the weight gain patterns observed in Figure 6A and C are specific to the Dp16 genotype. Dp16 females gained substantially more weight during the first six weeks of HFD before WT females caught up. In contrast, Dp16 males showed robust initial weight gain comparable to WT controls, but their weight plateaued after seven weeks while WT controls continued to gain, leading to a clear divergence (Fig. 6A).

      Other metabolic parameters also appear specific to the Dp16 model. On a standard chow diet, WT mice of both sexes generally do not exhibit glucose intolerance, insulin resistance, dysregulated lipoprotein profiles (VLDL-TG), or an impaired capacity to handle lipid loads. We observed all of these features in our Dp16 male and female mice (Fig. 3). Furthermore, transcriptomic analyses of Dp16 mice on standard chow revealed gene signatures of inflammation, fibrosis, and oxidative stress that are absent in WT mice.

      When challenged with HFD, while WT mice typically develop glucose intolerance and insulin resistance, the triplicated genes in Dp16 mice significantly exacerbated this metabolic deterioration. This is reflected in the worsening of glucose control and insulin sensitivity observed in our tolerance tests.

      In summary, most of these metabolic features are specific to Dp16 mice on a standard chow diet and are further exacerbated when combined with a high-fat diet.

      (3) Food intake data is difficult to interpret when weight has already diverged, as bigger animals will eat more food. Hence, the higher food may be a consequence rather than a cause of the weight gain (data in Figure 1).

      The reviewer makes a valid point. Since physical activity and energy expenditure do not differ significantly between Dp16 females and WT controls (Fig. 2F), the observed increase in food intake may indeed contribute to the higher body weights in Dp16 female mice.

      To rigorously confirm this, food intake would need to be measured between 6 and 8 weeks of age, prior to the divergence in body weight. Unfortunately, we did not measure food intake at that earlier time point.

      (4) The n numbers seem to vary significantly. For example, the use of n=6 for metabolic studies is generally rather small and underpowered. For the seahorse data, another concern is the snap freezing of samples before Seahorse assessment. For example, snap freezing of samples has been shown to increase certain metabolites. Freeze-thaw tissues often show a significant reduction in optical redox ratio.

      Regarding the transcriptomics and metabolomics studies, we utilized N=6 mice per tissue per sex. While we agree that a larger sample size is always preferable, the high cost of OMICS analyses covering 144 RNA-seq and 48 metabolomics samples limited our capacity to increase this number. However, N=6 remains a robust and standard approach for these specific assays. For the majority of our other in vivo and ex vivo data, we employed a higher sample size of 12-15 mice per genotype per sex to ensure statistical rigour. For a few assays, we have sample size of over 20.

      Regarding the respirometry analysis, we acknowledge the limitations of using frozen tissue. We chose this method because it allowed us to perform Seahorse assays on multiple tissues from 9-10 mice, which is a significant sample size for this type of analysis. The alternative isolating mitochondria from fresh tissue would have restricted our ability to process multiple tissues from a large number of animals on the same day due to the length of the protocol. We believe this trade-off was necessary to maintain a high sample size across various tissues.

      (5) For oestradiol measurements, were the samples taken at the same times within the estrous cycle? This may affect the comparability of female Dp16 and WT mice?

      Regarding our protocol, blood samples were collected between 11:00 AM and noon, with food removed two hours prior. While we did not specifically monitor the oestrous cycle of the female mice, serum samples for both the Dp16 females and WT littermates were collected on the same day and at the same time to ensure comparability across the groups.

      (6) Body weight reduction and organ size reduction on an HFD are especially interesting. Could enhanced inflammation and fibrosis be the root cause of this? Are there other mouse models where this is the reason?

      On a high-fat diet, we observed a reduction in iWAT and gWAT fat depot weights in both male and female Dp16 mice, which is consistent with their lower overall body weights (Fig. 6 - figure supplement 3). Conversely, Dp16 females fed a high-fat diet showed increased heart and kidney weights. Despite their lower adiposity, the Dp16 mice on this diet exhibited greater insulin resistance and glucose intolerance (Fig. 7). This suggests that the worsening of glucose control is independent of obesity. While we observed signatures of inflammation and fibrosis, we do not yet have direct mechanistic evidence demonstrating that these factors causally impaired glucose and lipid metabolism.

      (7) The authors are circumspect throughout to avoid over-claiming, as the majority of data is observational. One exception: "Many bile acids serve as ligands for nuclear hormone receptors (e.g., FRX and TGR5) that control various aspects of glucose and lipid metabolism (74, 75), and extensive changes in circulating bile acids are contributing, at least in part, to the systemic metabolic phenotypes in Dp16 mice." The authors have not shown a direct link between bile acids and metabolism in this model. Please edit.

      We have edited the text accordingly.

      Minor:

      (1)"Most human studies at the whole-body level are limited to assessing the impact of trisomy 21 on food intake, adiposity, physical activity level, and energy expenditure in adolescents or adults with DS"

      While we were uncertain of the reviewer's specific intent regarding the suggested changes, we have rephrased the sentence for clarity.

      (2) It is somewhat surprising that T3 is elevated, although there are reports of T3 elevation in visceral obesity in humans (e.g., Sun Nam et al., Obes Res Clin Pract, 2010).

      We observed that T3 levels did not differ by genotype in mice of either sex when fed a standard chow (Fig. 2 - figure supplement 5). However, we noted elevated T3 levels in both male and female Dp16 mice on a high-fat diet (Fig. 6 - figure supplement 2). While increased T3 levels correlated with higher physical activity and a modest increase in metabolic rate in Dp16 females, this was not observed in males (Fig. 6). We do not currently have a clear explanation for these findings. Given that individuals with Down syndrome often present with hypothyroidism and lower T3 levels, this discrepancy may reflect a species-specific difference between humans and mice.

      (3) Please can the authors clarify the percentage gene coverage, as this is quoted as ~58% of Hsa21 gene orthologs or ~65% of the Hsa21 gene orthologs, where the same reference is used.

      We apologize for the confusion. The number of triplicated genes in Dp16 mice corresponds to ~58% of Hsa21 genes (PMID: 26765563). We have corrected the typographical error in the text.

      (4) "segmental duplication model carrying a majority of the triplicated Hsa21 gene orthologs" for this given percentage majority sounds too strong, and the use of percentage is recommended.

      We have modified the text accordingly.

      (5) It is puzzling that in female gWAT with 7 triplicated Hsa21 gene orthologs (Rbm11, Chodl, Cldn8, Sh3bgr, Igsf5, Itgb2l, and Tmprss2). Could this be a technical issue? Was the reduced expression quantified by RT-Q-PCR?

      We have examined the normalized counts in the RNA-seq data for the seven genes in question, and the results do not appear to be an artifact. The sample size for this data is six mice per tissue per sex. In general, we prefer utilizing raw and normalized counts from RNA sequencing because there is a linear relationship between transcript amount and raw counts that is independent of housekeeping genes. In contrast, RT-qPCR involves mRNA amplification and requires expression to be normalized by one or more housekeeping genes (such as GAPDH, β-actin, 36B4, or ubiquitin) under the assumption that their levels remain constant.

      (6) The difference in body temperature is of interest. In male Dp16 mice, there is an increase in core temperature and a lowering of body temperature in females. In female Dp16 mice, higher estradiol levels have been stated by the authors to contribute to lower body temperature and higher physical activity (69-72). I am uncertain if the references are all relevant, as some relate to ovariectomized animals. No explanation is given for males.

      We currently do not have an explanation for why Dp16 males on a chow diet exhibit higher core body temperature, while Dp16 females show lower body temperatures. Although elevated T3 levels can increase body temperature, we have ruled this out; our data indicates there are no significant differences in T3 levels between genotypes for either sex on a chow diet.

      (7) The authors find a higher percentage heart weight in Dp16 mice on HFD and comment in the discussion that this is in keeping with "high-fat diet-induced cardiac hypertrophy". From what I can see, no histology has been performed to justify this statement. Furthermore, it would be useful to understand which animals had congenital heart disease in the first instance.

      We have modified the text accordingly. Unfortunately, we do not have histology data on the heart to inform us on whether some of our mice had congenital heart disease.

      Reviewer #2 (Recommendations for the authors):

      (1) The authors should comment on the dosing method of glucose/insulin/lipid in the tolerance tests to acknowledge that differences in body weight may affect these tests. In addition, I encourage the authors to present ITT data as raw data, and not % change.

      In response to the reviewer’s comments, we have updated the ITT data plots to show raw data rather than percentage change. Regarding the dosing methodology, we maintain that basing dosage on body weight is appropriate. This approach is consistent with the vast majority of published rodent studies, as blood volume and metabolic tissues—such as skeletal muscle and adipose tissue—scale with body weight. Standardizing dose independently of body weight could lead to erroneous conclusions.

      (2) It would be useful for the authors to include a discussion on the likely specific tissue involvement in the whole-body metabolic disturbance. From my reading of the manuscript, there seems to be data suggesting functional and transcriptional dysfunction across most tissues, but do the authors suggest there is a dominant tissue in this regard?

      Due to the triplication of large number of genes on human chromosome 21, people with Down syndrome exhibit deficits across most organ systems (PMID: 32029743). Metabolic homeostasis also involves multiple tissues and cell types (adipose tissues, liver, skeletal muscle, pancreas, gut, hypothalamus, and immune cells). Most of the triplicated genes do express across these tissues. Our data indicate metabolic dysregulation across adipose tissues (white and brown), liver, skeletal muscle, and hypothalamus. Given the complex genetic perturbations of the Down syndrome mouse model, we do not think that there is a dominant tissue that contributes disproportionately to the systemic metabolic dysfunction phenotypes we observed in the Dp16 mice. Rather, we think that the metabolic phenotype is due to the combined deficits across multiple organs and tissues. As we do not have data to support the disproportionate contribution of any one tissue, we therefore did not speculate on the dominant contribution of any single tissue in the Discussion.

      (3) Related to this, muscle lipid is thought to be a major driver of muscle insulin resistance. Do the authors have measures of muscle lipid accumulation? This might be particularly interesting in the HFD models.

      Unfortunately, we did not measure lipid content in the skeletal muscle during this study. For the chow-fed mice, the entire gastrocnemius muscle was used for RNA isolation to perform RNA sequencing, and no tissue remains for additional analysis. Regarding the HFD-fed group, skeletal muscle was not collected at the termination of the study. As a result, we are unable to provide the requested lipid analysis data.

      (4) For mitochondrial analyses - do the authors have measures of total tissue mitochondria, and might changes in mitochondria abundance be driving some of these differences?

      For all our mitochondrial respiration analyses, we normalized the data to mitochondrial content as quantified by the MTDR assay (PMID: 32432379; PMID: 39704485). These results indicate that for a given amount of mitochondrial content, respiration as measured by the Seahorse assay is reduced in Dp16 mouse tissues, specifically in the BAT and liver.

      (5) To broaden the scope and interest, can the authors compare the transcriptional or metabolomic data to what has been found in non-DS insulin resistance (humans or mice), for example? This may help to highlight the key changes in metabolism that are causal for specific phenotypes.

      Overall, this is a comprehensive assessment of metabolism in a DS model.

      We appreciate the reviewer’s suggestion. However, given the vast number of published datasets on non-DS insulin resistance in both humans and mice, comparisons would yield varying results depending on the specific datasets selected. Consequently, we feel that such an analysis is beyond the scope of this study. We would like to highlight that many of the processes dysregulated in Dp16 mice as identified through our pan-tissue transcriptomes and metabolomes align with those frequently observed in non-DS insulin resistance. These include signatures of chronic low-grade inflammation, fibrosis, ER and oxidative stress, and impaired glucose and lipid metabolism.

      Reviewer #3 (Recommendations for the authors):

      It is slightly disconcerting that Figure 5 - Figure Supplements 2-5 are referred to in the text before the data in Figure 5 are discussed. It might make sense to indicate that the data are discussed further below (assuming that the authors do not wish to renumber these figures).

      We have fixed this issue raised by the reviewer.

    1. eLife Assessment

      This study presents an important large-scale behavioral and transcriptomic analysis of Drosophila that are heterozygous for putative loss-of-function alleles of homologs of human genes that have been linked to autism spectrum disorders. The authors consider 48 genes as hits from their screen, which show significant behavioral alterations in sleep, basal activity, and/or social behavior, and significant sexual dimorphism. The authors then focus on the domino/SRCAP gene as a candidate regulator of sleep, social behavior, transcriptional programs, and RNA splicing. The work generates a solid dataset and applies quantitative analytical approaches that will be of interest to researchers in the field, yet the evidence presented remains incomplete because issues of genetic background need to be further addressed.

    2. Joint public review:

      Summary:

      In this study, Stirtz et al., performed a targeted screen of 80 Drosophila strains carrying heterozygous MiMIC insertions in genes that are homologous to human genes that have been linked to autism spectrum disorders (ASD). This is an important and timely topic, as human genetic studies have identified a large number of ASD risk genes, yet the functional characterization of many of these candidates remains limited. The authors identify 48 putative mutants with altered sleep, activity, or social behavior. They then focus on one hit, domino (the orthologue of human SRCAP), for which the heterozygous MiMIC mutants show altered behavior in males but not in females. They show that domino is a candidate regulator of sleep, activity, social behavior, transcriptional programs, and RNA splicing. The authors molecularly validate that the heterozygous MiMIC insertion in domino causes a 50% reduction in gene expression, and use RNA-seq to show that the heterozygous MiMIC males and females have altered gene expression profiles and splicing patterns. Finally, they use immunostaining against the commonly used synaptic marker, Bruchpilot, to show that both males and female heterozygous domino flies express a higher immunosignal compared to the wild-type control.

      Strengths:

      This work provides potential genetic links between human ASD genes and fly behavioral phenotypes. Overall, it represents an ambitious and technically valuable effort that generates a substantial behavioral dataset across a large number of ASD-associated orthologues and develops quantitative analytical approaches to extract information from complex phenotypes. One strength of this study is its focus on heterozygous mutants, which is more representative of human scenarios. The study also provides a potentially useful resource for the field, particularly through the identification of candidate genes and behavioral signatures that may warrant future mechanistic investigations. The screening experiments and analysis are well conceived, the manuscript is very clearly written and is easily understandable, and the concise, accurate interpretations for each result, aided by clear graphic representation of multiple dimensions in the behaviors tested, allow the reader to understand the paper with ease.

      Weaknesses:

      The work presents a few important weaknesses, especially with regard to the genetic and molecular validation of the mutants identified.

      (1) The authors validate that the MiMIC insertion affects the gene of interest only for the domino gene. The original MiMIC study (PMID: 25824290, eLife) reported that ~8% (5/63) MiMIC lines do not function as strong loss-of-function alleles. Thus, of the 48 hits identified here, one would estimate that ~4 of them may not cause the loss of function of the gene defined by the MiMIC insertion. To strengthen their claim, the authors would need to confirm that all of the MiMIC lines that they consider as hits do indeed significantly reduce the expression of the target genes.

      (2) Although the authors document that they validated the phenotype seen in the domino MiMIC line using a second mutant allele (Trojan), these two mutants share the same genetic background because the Trojan line was made from the MiMIC line via recombinase-mediated cassette exchange. Thus, the phenotype seen in the MiMIC and Trojan lines would need to be confirmed using a completely independent mutant in order to demonstrate that the reported behavioral, molecular, and synaptic defects reported can be fully attributed to the partial loss of domino function. Also, while the authors performed an RNA-seq experiment in both the MiMIC and Trojan lines, they do not show whether the Bruchpilot phenotype is also seen in the Trojan allele. Thus, this phenotype would also need to be examined in the Trojan allele or, preferably, in a mutant allele that is independent of the MiMIC line.

      (3) The RNA-seq results would benefit from a discussion of potential compensatory or secondary transcriptional effects resulting from the constitutive domino reduction, particularly since the expected global bias toward transcriptional downregulation was not observed. In addition, some neurobiological interpretations appear stronger than currently justified by the literature or the data presented, particularly regarding the Bruchpilot immunoreactivity analyses and their relationship to sleep-regulatory circuits. Additional validation using better-established sleep-related neuronal populations, together with a clearer discussion of sex-specific effects and alternative interpretations of the observed phenotypes, would substantially strengthen the manuscript.

      (4) An explanation of the extensive PCA analyses performed would help the naïve reader.

    1. eLife Assessment

      In this useful Tools & Resources article, the authors describe a new cryogenic light microscopy design and characterize its temperature and spatial stability. This compelling system avoids the challenges associated with vacuum-based designs, particularly vacuum transfer systems, which are difficult to engineer. A key advantage of the system is that it reduces ice contamination and drift, which are the primary challenges in open cryostat systems.

    2. Reviewer #1 (Public review):

      Summary:

      In the manuscript "A stable cryogenic fluorescence microscope for correlative super-resolution light and electron microscopy," the authors demonstrate a new cryogenic light microscopy design and characterize its temperature and spatial stability. The manuscript does a good job of reviewing the state of the field and highlights the need for improved cryogenic microscope stages. The system avoids challenges associated with vacuum-based designs, particularly vacuum transfer systems that can be difficult to engineer, while also showing minimal ice contamination and drift, which are the primary challenges associated with open cryostat systems.

      Strengths:

      The key strengths of the manuscript are the simple design and the significant level of detail provided in the description of the cryogenic stage. This represents a valuable step forward for the field by providing a home-built, non-vacuum stage design that others can emulate.

      Weaknesses:

      There are only minor weaknesses or issues to address, which, if resolved, would strengthen the manuscript overall.

      (1) A key element of the design gets little attention, which is the plastic cap for the objective. It is not entirely clear to the reader how this is being used except as something of a thermal break between the cryogen environment and the objective, but there are some questions. Is the objective housing touching the plastic cap? Where is the front of the cap relative to the front objective lens? Is the front objective lens exposed to the cryogenic environment? Could the authors provide some 3D views of that in an SI figure? This would help clarify.

      (2) The refilling system is not shown in the diagrams provided in Figure 1 and S1 in sufficient detail. How is the system mechanically coupled to the dewar on the microscope stage? Are there any concerns about coupling vibrations onto the table?

      (3) There is a description on page 6 that a rectangular aperture is used to align the excitation with the position and orientation of the sample. I know the authors are using this for excitation of the lamella, but without saying so in this text, it is confusing. I would consider stating that this is for future work involving excitation of lamella and then citing their preprint.

      (4) In Figure 2d, the z-drift is shown with the focus lock correction applied. This is highly relevant, but I also think it would be good to plot the z position plus the stage position in an SI figure. This will give a better idea of the mechanical stability of the system. Also, in this figure, I wonder if the authors could comment on the source of the jumps in lateral position. For example, just before 30 minutes. Lastly, I would make the lower plot have a tighter y-axis range. It is hard to see anything, hence the inset.

      (5) The ice contamination looks minimal in Figure 3. I think it would benefit the manuscript to have lower magnification images as well, to show the level of ice contamination across a representative square. This would be good, but only if the authors have it in hand.

      (6) In Figure 4b, the y-axis is unclear. It looks like it has been normalized. Consider revising.

      (7) A fluorescence intensity trace for the data shown in Figures 4c and f would be helpful to show the single-molecule behavior.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript reports the development of a cryo super-resolution fluorescence microscopy system. The authors demonstrate that they can achieve a mechanical and thermal stability that is sufficient to perform cryo-SMLM over the course of several hours. Focus instability is compensated for by tracking a fluorescent bead for its movement in the axial direction and adjusting the sample stage accordingly during data acquisition. Lateral instabilities are corrected after data acquisition. An enclosure around the microscope allows to significantly reduce ice contamination during cryo-SMLM imaging and sample transfer. The authors show an example of correlative cryo-SMLM and cryo-ET imaging achieved with their microscope system, which depicts the distribution of FtsZ-rsEGFP2 in E. coli.

      Strengths:

      The authors have designed a microscopy system for SR-cryo-CLEM, which achieves high stability while reducing complexity and costs substantially when compared to vacuum-insulated systems (e.g., Hoffman et al., 2020). They also provide software for controlling the microscope and data acquisition. This lowers the barrier for other labs to implement SR-cryo-CLEM into existing cryo-ET workflows. Reduction of ice contamination helps to increase throughput, which is currently one of the biggest bottlenecks for SR-cryo-CLEM.

      Weaknesses:

      To correct for focus drift, the authors track a fluorescent bead in the far-red channel. This is possible for bacterial samples as used in this work, as beads can easily be introduced to surround the cells.

      Recommendations:

      (1) It is not discussed how this can be achieved in other samples than bacterial samples, such as lamellae in mammalian cells. Here, it would be much more difficult to introduce bright point-like markers with far-red fluorescence that would be distributed in the entire cell to capture at least one in the final lamella. Furthermore, it might be important to know for readers whether the far-red channel has to be sacrificed entirely for the focus correction.

      (2) The authors show an application of SR-cryo-CLEM imaging of FtsZ-rsEGFP2 in E. coli. In the chosen correlative example (Figure 4d.f), no clear structure can be seen in the fluorescent images. The overview image (Figure 4d) shows no distinct signal in the cell, as it is shown for the non-correlative example in Figure 4a. The cryo-SMLM image (Figure 4f) does not show any ring-like features or accumulations of signals at the constriction site, as would be expected for a projecting along the optical axis. A clearer application example, which would show how increased resolution in cryo fluorescence microscopy enables resolving certain structural details or adds information not accessible in cryo electron tomography, would have strengthened the work. Particularly if taking into consideration that bacteria have a strong auto-fluorescence in the green range (Dahlberg et al., 2020), which could lead to high background or false positive localizations when using green fluorophores as labels.

      (3) Access to CAD drawings (particularly for custom-made parts, such as cryostat or humidity enclosure) and a parts list is highly important for other researchers who would like to set up this SR-cryo-CLEM system in their own lab or institution. This is currently missing and, therefore, creating a hurdle for a wider adaptation of the technique.

    4. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In the manuscript "A stable cryogenic fluorescence microscope for correlative super-resolution light and electron microscopy," the authors demonstrate a new cryogenic light microscopy design and characterize its temperature and spatial stability. The manuscript does a good job of reviewing the state of the field and highlights the need for improved cryogenic microscope stages. The system avoids challenges associated with vacuum-based designs, particularly vacuum transfer systems that can be difficult to engineer, while also showing minimal ice contamination and drift, which are the primary challenges associated with open cryostat systems.

      Strengths:

      The key strengths of the manuscript are the simple design and the significant level of detail provided in the description of the cryogenic stage. This represents a valuable step forward for the field by providing a home-built, non-vacuum stage design that others can emulate.

      We thank the reviewer for their positive assessment and strive to address the weaknesses they have constructively raised below.

      Weaknesses:

      There are only minor weaknesses or issues to address, which, if resolved, would strengthen the manuscript overall.

      (1) A key element of the design gets little attention, which is the plastic cap for the objective. It is not entirely clear to the reader how this is being used except as something of a thermal break between the cryogen environment and the objective, but there are some questions. Is the objective housing touching the plastic cap? Where is the front of the cap relative to the front objective lens? Is the front objective lens exposed to the cryogenic environment? Could the authors provide some 3D views of that in an SI figure? This would help clarify.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we will include a new supplementary figure (Fig. S2) providing detailed 3D views of the copper adapter, microscope objective, and plastic cap. The figure will show that the rim surrounding the front lens of the objective is covered by the plastic cap to provide thermal insulation between the objective housing and the cryogenic environment (Fig. S2b). We will also clarify that the front surface of the cap is levelled with the front objective lens to maintain the full working distance of the objective while allowing for axial movement of the z-stage. Finally, we will explicitly state that the front objective lens is exposed to the cryogenic environment (cold nitrogen gas).

      (2) The refilling system is not shown in the diagrams provided in Figure 1 and S1 in sufficient detail. How is the system mechanically coupled to the dewar on the microscope stage? Are there any concerns about coupling vibrations onto the table?

      To minimise vibrations arising from the nitrogen refilling pumps, the cryostat and liquid nitrogen tubing are mechanically decoupled from the microscope cage system, objective, translation stages, and sample. Specifically, the cryostat and nitrogen tubing are supported independently on a laboratory jack and surround the cage system without rigid mechanical contact. In the revised manuscript, we will update Fig. S1a, b to illustrate the liquid nitrogen tubing and refilling system more clearly. In addition, we will include a new supplementary figure (Fig. S3) to show the detailed cryostat design, refilling tubing, and temperature sensor position.

      (3) There is a description on page 6 that a rectangular aperture is used to align the excitation with the position and orientation of the sample. I know the authors are using this for excitation of the lamella, but without saying so in this text, it is confusing. I would consider stating that this is for future work involving excitation of lamella and then citing their preprint.

      We agree that the purpose of the rectangular aperture should be made clearer. In line with the suggestion from the reviewer, in the revised manuscript, we will briefly explain that the aperture is intended for selective illumination, such as in applications to cryo-FIB lamellae, and will cite our recent preprint describing this approach.

      (4) In Figure 2d, the z-drift is shown with the focus lock correction applied. This is highly relevant, but I also think it would be good to plot the z position plus the stage position in an SI figure. This will give a better idea of the mechanical stability of the system. Also, in this figure, I wonder if the authors could comment on the source of the jumps in lateral position. For example, just before 30 minutes. Lastly, I would make the lower plot have a tighter y-axis range. It is hard to see anything, hence the inset.

      We thank the reviewer for this suggestion. In the revised manuscript, we will include an additional supplementary figure (Fig. S4) showing the axial drift measured without focus-lock correction to illustrate the intrinsic mechanical stability of the microscope. We will also clarify that the periodic lateral displacement observed along the x-direction (approximately 300 nm amplitude with a period of ~22 minutes) arises from slight lateral repositioning accompanying z-stage stepping during focus-lock operation, likely due to mechanical coupling between the axes of the translation stage. We will revise the lower panel of Fig. 2d by reducing the y-axis range to improve data visibility.

      (5) The ice contamination looks minimal in Figure 3. I think it would benefit the manuscript to have lower magnification images as well, to show the level of ice contamination across a representative square. This would be good, but only if the authors have it in hand.

      We agree that this would be useful. In the revised manuscript, we will update Fig. 3 to include two additional low and intermediate-magnification cryo-EM images showing a representative grid square and a zoomed-in region of it, including a few grid holes. These images provide an overview of the ice contamination across a substantially larger field of view.

      (6) In Figure 4b, the y-axis is unclear. It looks like it has been normalized. Consider revising.

      The y-axis in Fig. 4b represents the localization rate (number of detected localizations per frame) within the selected ROI in Fig.4c and was not normalized. The values were calculated in SMAP by binning the localization frames into 100 temporal bins and dividing the number of localizations in each bin by the corresponding bin width, resulting in units of localizations per frame. Therefore, values close to 1 indicate approximately one localization detected per frame at that time point. To avoid potential confusion regarding the interpretation of this representation, we will replace this plot in the revised manuscript with a more explicit visualization showing the number of detected localizations per defined number of frames as a function of time (frame number) for the specific ROI shown in Fig. 4c.

      (7) A fluorescence intensity trace for the data shown in Figures 4c and f would be helpful to show the single-molecule behavior.

      In the revised manuscript, we will add fluorescence intensity traces corresponding to the single-molecule events shown in Fig. 4c and Fig. 4f to further demonstrate their single-molecule emission characteristics.

      Reviewer #2 (Public review):

      Summary:

      This manuscript reports the development of a cryo super-resolution fluorescence microscopy system. The authors demonstrate that they can achieve a mechanical and thermal stability that is sufficient to perform cryo-SMLM over the course of several hours. Focus instability is compensated for by tracking a fluorescent bead for its movement in the axial direction and adjusting the sample stage accordingly during data acquisition. Lateral instabilities are corrected after data acquisition. An enclosure around the microscope allows to significantly reduce ice contamination during cryo-SMLM imaging and sample transfer. The authors show an example of correlative cryo-SMLM and cryo-ET imaging achieved with their microscope system, which depicts the distribution of FtsZ-rsEGFP2 in E. coli.

      Strengths:

      The authors have designed a microscopy system for SR-cryo-CLEM, which achieves high stability while reducing complexity and costs substantially when compared to vacuum-insulated systems (e.g., Hoffman et al., 2020). They also provide software for controlling the microscope and data acquisition. This lowers the barrier for other labs to implement SR-cryo-CLEM into existing cryo-ET workflows. Reduction of ice contamination helps to increase throughput, which is currently one of the biggest bottlenecks for SR-cryo-CLEM.

      We thank the reviewer for their critical assessment, and for their suggestions below which we have used to improve the manuscript.

      Weaknesses:

      To correct for focus drift, the authors track a fluorescent bead in the far-red channel. This is possible for bacterial samples as used in this work, as beads can easily be introduced to surround the cells.

      Recommendations:

      (1) It is not discussed how this can be achieved in other samples than bacterial samples, such as lamellae in mammalian cells. Here, it would be much more difficult to introduce bright point-like markers with far-red fluorescence that would be distributed in the entire cell to capture at least one in the final lamella. Furthermore, it might be important to know for readers whether the far-red channel has to be sacrificed entirely for the focus correction.

      We thank the reviewer for highlighting this point. We agree that focus stabilization strategies for cryo-FIB lamellae are likely to differ from those used for the individual bacterial cell samples. For lateral drift correction, the presence of a single continuously detectable bright feature within the field of view is sufficient. Importantly, this feature does not need to be a fluorescent bead; any stable signal that can be continuously detected by the camera can serve as a suitable reference for drift correction. We will expand the Discussion to describe potential strategies for stable cryo-SMLM imaging, including the use of intrinsic sample or lamella features for autofocus, minimal fiducial-based approaches, and the practical implications of dedicating the far-red channel to focus stabilization.

      Furthermore, in the revised manuscript, we will include a new supplementary figure (Fig. S4) demonstrating the intrinsic axial stability of the microscope in the absence of active focus-lock correction. These measurements show that the system remains within the objective's depth of focus for a relatively long time, providing adequate stability for experiments in which far-red fluorescent fiducial beads are unavailable, such as cryo-FIB lamella imaging.

      (2) The authors show an application of SR-cryo-CLEM imaging of FtsZ-rsEGFP2 in E. coli. In the chosen correlative example (Figure 4d.f), no clear structure can be seen in the fluorescent images. The overview image (Figure 4d) shows no distinct signal in the cell, as it is shown for the non-correlative example in Figure 4a. The cryo-SMLM image (Figure 4f) does not show any ring-like features or accumulations of signals at the constriction site, as would be expected for a projecting along the optical axis. A clearer application example, which would show how increased resolution in cryo fluorescence microscopy enables resolving certain structural details or adds information not accessible in cryo electron tomography, would have strengthened the work. Particularly if taking into consideration that bacteria have a strong auto-fluorescence in the green range (Dahlberg et al., 2020), which could lead to high background or false positive localizations when using green fluorophores as labels.

      We thank the reviewer for this thoughtful comment. We agree that a correlative example displaying more pronounced structural features would further illustrate the capabilities of cryo-SMLM. However, the primary aim of the present work is the development and characterization of a robust cryogenic super-resolution microscope for reliable cryo-SMLM and correlative cryo-CLEM, rather than the demonstration of new biological applications. The utility of correlative cryo-SMLM/cryo-ET for resolving cellular structures has already been established in previous studies, including those employing rsEGFP2-labelled targets.

      The correlative dataset presented here is intended to demonstrate the compatibility of the microscope with cryo-CLEM workflows rather than to provide detailed biological insight. Moreover, the use of intact E. coli cells imposes inherent limitations on the ultrastructural information accessible by cryo-electron tomography; overcoming these limitations would typically require specimen thinning, for example, by cryo-focused ion beam (cryo-FIB) milling, which is beyond the scope of the present work.

      Regarding the concern about auto-fluorescence, elevated background fluorescence is not unique to bacterial samples or green fluorescent proteins but is a general consideration in cryo-SMLM that depends on the specimen and imaging conditions. While auto-fluorescence may reduce image contrast, it does not affect the conclusions of this work, which focuses on the design and performance of the microscope.

      (3) Access to CAD drawings (particularly for custom-made parts, such as cryostat or humidity enclosure) and a parts list is highly important for other researchers who would like to set up this SR-cryo-CLEM system in their own lab or institution. This is currently missing and, therefore, creating a hurdle for a wider adaptation of the technique.

      Thank you for this useful suggestion. In the revised manuscript, we will make available the complete SolidWorks CAD files for all custom-designed components, together with a comprehensive parts list and the full assembly corresponding to Fig. S1 as supplementary materials.

    1. eLife Assessment

      This manuscript describes a valuable study of the mechanism by which acetylation on the histone H3 core domain regulates RNA polymerase II transcription passing through nucleosomes. The authors provide convincing evidence that acetylation influences transcription in a context-specific fashion. Some questions relating to the static nucleosome structures and the polymerase passage remain, but this manuscript will be of considerable interest to researchers in the chromatin and transcription fields.

    2. Reviewer #1 (Public review):

      Summary:

      The authors investigate how site-specific acetylation within the histone H3 folded domain affects RNA polymerase II transcription through nucleosomes. They focus on H3K56ac, H3K64ac, and H3K122ac, prepare chemically defined nucleosomes carrying each modification, and compare their effects using an in vitro transcription assay, cryo-electron microscopy structures, and micrococcal nuclease sensitivity assays.

      The main finding is that H3K56ac and H3K122ac increase production of full-length run-off transcripts and reduce pausing near the nucleosomal dyad region, whereas H3K64ac has little detectable effect under the same reconstituted conditions. The structural analyses suggest that H3K56ac weakens or destabilizes DNA near the entry/exit region, while H3K122ac alters histone-DNA contacts near the dyad. These observations support a model in which different acetylation sites within the H3 folded domain influence nucleosomal transcription barriers through distinct local effects on histone-DNA interactions.

      This is a useful study because it examines histone core-domain acetylation using chemically defined nucleosomes and directly compares several modifications in the same experimental system. However, the broader cellular context of these modifications is not sufficiently developed, and some mechanistic conclusions rely on correlations between static nucleosome structures and endpoint transcription assays rather than direct observation of polymerase passage through modified nucleosomes.

      Strengths:

      (1) The study uses site-specifically acetylated H3 proteins and reconstituted nucleosomes, allowing direct comparison of H3K56ac, H3K64ac, and H3K122ac under controlled conditions.

      (2) The combination of transcription assays, cryo-electron microscopy, and nuclease sensitivity assays provides multiple lines of evidence, particularly for increased DNA end flexibility in H3K56ac nucleosomes.

      (3) The authors analyze unmodified, H3K56ac, H3K64ac, and H3K122ac nucleosomes in parallel, with reported structural resolutions of approximately 3 Angstroms and accompanying validation materials.

      (4) The negative result for H3K64ac is informative, because it distinguishes the direct effect of this modification in a minimal reconstituted system from prior cellular associations with active chromatin and histone eviction.<br /> The comparison with H3 N-terminal acetylation highlights that acetylation within the folded domain may affect transcription at different positions or by different mechanisms than tail acetylation.

      Weaknesses:

      The rationale for focusing on H3K56ac, H3K64ac, and H3K122ac has not been developed sufficiently. The manuscript would benefit from a clearer summary of what is known about the abundance of these modifications in cells, the enzymes or histone metabolic pathways that may introduce or remove them, and whether they are thought to occur before histone deposition, on assembled nucleosomes, or during nucleosome remodeling.

      The central mechanistic model is based mainly on correlations between structures of free nucleosomes and endpoint transcription assays. The study does not directly observe RNA polymerase II paused at or passing through the relevant nucleosomal positions, so the proposed link between local structural changes and reduced pausing should be stated with appropriate caution.

      The H3K56ac interpretation is supported by both structural observations and nuclease sensitivity data, but the map comparison underlying the reduced entry/exit DNA density is still mostly qualitative. The manuscript should more clearly state the map comparison conditions, such as contouring and local map quality, so that non-specialist readers can judge how robust the local density differences are.

      The H3K122ac mechanism is plausible, but the evidence for dyad destabilization is more indirect. The main support comes from the orientation of the K122 side chain and its distance from DNA, while an independent biochemical test of dyad-region destabilization is not provided.

      The transcription assay appears to include statistical testing, but the figure legend and methods should more clearly state which tests were used, what comparisons were made, how n was defined, and whether multiple-comparison correction was applied.

      The relationship between the 198 bp transcription template, the linker DNA, the 9-base mismatched region, and the DNA regions modeled in the cryo-electron microscopy structures is somewhat difficult to follow. This does not necessarily require new experiments, but a clearer explanation would help readers connect the transcription assay design with the structural models.

      The use of H3.2 C110A for chemical ligation and the use of the PL2-6 single-chain antibody fragment for cryo-electron microscopy sample stabilization are reasonable technical choices, but their purposes and possible effects on interpretation should be explained more clearly for readers outside structural biology.

      Because the work uses a minimal in vitro system with human nucleosomes and Komagataella phaffii RNA polymerase II/TFIIS, the conclusions should be limited to direct physical effects on nucleosome transcription barriers unless cellular cofactors, remodelers, histone chaperones, additional modifications, and nucleosome positioning are addressed or discussed.

    3. Reviewer #2 (Public review):

      Summary:

      Chromatin regulates a wide range of biological processes. The nucleosome, composed of 147 bp of DNA wrapped around a histone octamer containing histones H2A, H2B, H3, and H4, is the fundamental unit of chromatin. Post-translational modifications of histone proteins regulate the dynamic properties of nucleosomes and thereby influence chromatin accessibility and gene expression. Among these modifications, lysine acetylation on histone H3 is closely associated with transcriptional activation. While the epigenetic functions of acetylation on the histone H3 N-terminal tail have been extensively studied, the molecular mechanisms by which acetylation within the histone H3 core domain, particularly at Lys56, Lys64, and Lys122, modulates nucleosome architecture to facilitate RNA polymerase II (RNAPII) transcription remain unclear.

      In this study, Oishi et al. investigated the effects of histone H3 acetylation at K56, K64, and K122 on RNAPII transcription using in vitro transcription assays. Furthermore, the authors determined the three-dimensional structures of nucleosomes containing these acetylation marks by cryo-electron microscopy single-particle analysis, revealing distinct structural dynamics depending on the acetylation site. Overall, this study advances our understanding of the molecular mechanisms linking histone H3 core acetylation to transcriptional regulation.

      Strengths:

      (1) Site-specifically acetylated histone H3 proteins were chemically synthesized using a unique and rational peptide ligation strategy, representing a major technical strength of this study.

      (2) The in vitro transcription assays demonstrated that H3K56ac and H3K122ac increase the production of run-off transcripts, whereas H3K64ac has little effect on transcription efficiency. These findings highlight the distinct functional roles of individual acetylation sites within the histone H3 core domain.

      (3) The cryo-EM structures of nucleosomes containing either H3K56ac or H3K122ac revealed that H3 acetylation weakens histone-DNA interactions, providing a structural basis for the observed effects on transcription.

      Weaknesses:

      (1) Although the biochemical and structural data are convincing and sufficiently support the authors' conclusions, complementary cellular experiments would further strengthen the physiological relevance of the in vitro findings. While such experiments are not essential for supporting the main claims of the study, they would enhance the overall impact and biological significance of the work.

      (2) Although the authors demonstrate the structural consequences of individual H3 core acetylation events, the study does not investigate potential synergistic effects among multiple acetylated lysine residues within the H3 core domain. Consequently, the relationship between combinatorial acetylation patterns and their collective impact on RNA polymerase II-mediated transcription remains unclear.

    4. Reviewer #3 (Public review):

      This is a short and punchy manuscript that nicely summarises the 4 structures that are determined and provides a basis for the differences seen for acetylation sites shown for RNAPII activity.

      The authors build on previous biochemical work that determined the functional outcomes of H3 core acetylation, adapting an assay they have previously used extensively to investigate RNAPII transcription on nucleosomes and, indeed, even H3 N-terminal tail acetylation. This assay is as such well set up and has a wealth of confirmatory previous studies from this lab and the authors are careful not to overanalyse their results, leading to robust and well-considered results. The structures are determined to a high resolution, allowing the interpretation put forward about side chain orientations, with clear densities shown for the regions of interest.

      Further discussion or experiments would strengthen the conclusions further:

      (1) The conclusion on the role of H3K56Acetylation could be strengthened, especially as the results are somewhat counterintuitive. It is conceptually surprising that acetylation near the entry/exit DNA that destabilises this region also leads to a reduced stall propensity at the dyad but has a limited effect at SHL5? While it can be explained by the clash at the dyad pause being reduced, the more direct effect of DNA breathing amplification would be expected to have a larger effect at SHL 5. Indeed, the density for DNA at SHL5 appears to be weaker in Figure 2A, suggesting the entry/exit DNA flexibility is amplified past this region.

      Perhaps another assay that looks more directly at the flexibility of the entry/exit DNA would be useful, either through restriction enzyme-mediated cleavage or FRET (DNA ends and H2AK119 labels), providing stronger evidence of this effect. MNase is rather indirect and similar to the RNAPII assay itself.

      Similarly, were the authors surprised by the modest effect (less than 2-fold) in transcriptional pause at SHL 0 for the K122Ac? Presumably, based on the model in Figure 4, this would be expected to be the area with the largest effect? The results of K56Ac and K122Ac almost seem swapped to what would be expected in Figure 1H. Further discussion of this observation would be useful.

      (2) Could the local weakening of DNA, especially at the dyad, be observed in the cryo-EM structures? Perhaps comparison of local resolution estimation differences in this region compared to unmodified would be useful.

      (3) Caution should be taken, and discussion should include that the structural data presented is after extensive processing. Many nucleosome averaging classes were discarded in the 3D classification steps (nicely summarised in Table 1 as "particles for 3d classification" and "particles in final map"). Indeed, it is likely that higher DNA flexibility particles would be thrown away during this processing step. This can be observed for K56Ac DNA ordering, for example, in Supplementary Figure S4, yellow and cyan classes from the round of 3D classification look to be high resolution and have a higher order of DNA, so there has been some selection here. How was this done? While this is not fully quantifiable, it gives an idea of the extent of wrapping. We would suggest discussing the methodological limitations and showing the models after the first auto refinement to see if the features discussed on end flexibility and dan ordering are retained.

      (4) Di Cerbo et al. (reference 13) showed acetylation at K64 alters salt stability and affects transcription. Why do the authors think there is a discrepancy, albeit with different assays? Direct reference and discussion of this in the text should be included.

      (5) Why was H3.2 used, while this is relatively abundant in mouse cells, human protein was used, and this appears to be less common than H3.1 and H3.3. We are sure that the effect is not likely to be substantive on structure (as shown by the Kurumizaka lab previously), but should be addressed in the text

    1. eLife Assessment

      This important study provides a mechanistic view of how antibody affinity maturation can reshape encounter-state landscapes and association pathways, with implications for understanding HIV antibody maturation and vaccine design. The results are solid, supported by a coherent integration of adaptive molecular dynamics, Markov state modeling, SPR kinetics, mutagenesis, and double-mutant cycle analysis, although aspects of the kinetic validation, MSM-state robustness, and causal interpretation would benefit from further support. The work will be of interest to immunologists, structural biologists, and computational biophysicists.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript uses simulations and MSMs paired with experimental binding assays to examine the binding mechanisms of different antibodies to their targets. The authors argue that contacts in encounter complexes play an important role in determining the association rates and binding affinities that distinguish more mature antibodies from less efficacious antibodies from earlier in the maturation process.

      Strengths:

      The idea is interesting, and the combination of computational models and experiments is a good direction.

      Weaknesses:

      The manuscript focuses heavily on kinetics, but it is not clear whether the simulations recapitulate the relative rates of binding of the two antibodies. The relationship between the simulated binding behavior and the experimentally observed kinetic differences is therefore not fully established.

      The comparison of committor probabilities or fluxes between the two antibodies may not be appropriate. These properties are related to the barrier height the system has to cross to move forward vs back to the starting state, under the simplifying assumption that the properties of other states aren't critical. Even in this simplified case, the same flux or committor probability could occur with very different barrier heights, e.g., rates or transition probabilities.

      Some claims are presented in a very qualitative way that people who aren't experts in MSMs may have difficulty tying to the results in Figure 1.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript addresses an important and underexplored question: how affinity maturation alters antibody encounter-state landscapes rather than simply improving bound-state affinity. The authors combine adaptive MD, Markov State Models (MSMs), transition path theory, mutagenesis, SPR kinetics, and double-mutant cycle analysis into a coherent story.

      Strengths:

      This manuscript presents a compelling computational and experimental analysis of antibody affinity maturation in the HIV-1 DH270 lineage. The main finding is that somatic mutations reshape encounter-state pathways through glycan-mediated steering rather than simply stabilizing the final bound state. This is novel and potentially important for vaccine design. The combination of adaptive MD, MSMs, SPR kinetics, and double-mutant cycle analysis is a major strength.

      Weaknesses:

      The proposed sequence that somatic mutations cause glycan capture, which causes reorientation, which causes enhanced association, is based on correlation rather than direct causality.

      The four MSM states are not convincingly explained, and the robustness of these states is unclear.

      The productive collision surface area analysis needs more quantitative data.

      The coupling energy values are near the uncertainty range. Some conclusions about long-range communication networks appear stronger than the data justify. The data support coupling, but they do not necessarily support detailed mechanistic networks.

      The study investigates one lineage, one epitope class, and one viral system. Hence, the generalization is limited.

    4. Reviewer #3 (Public review):

      Summary:

      In this work, the authors set out to characterise how encounter states between antibodies and antigens evolve during affinity maturation through molecular dynamics simulations and Markov state modeling. They demonstrate how early glycan-mediated interactions increased association rates rather than modifying the final bound state.

      Strengths:

      The computational approach is backed up by experimental results and allows for visualising otherwise too short-lived association states, thus allowing to discriminate between different lineages.

      Weaknesses:

      The figures and captions are not always clear about what they are trying to show. The choice of CVs is not sufficiently discussed.

    1. eLife Assessment

      Muetter et al. provide an important argument that luminescence is a reliable, high-throughput alternative to colony-forming units (CFU) for super-MIC investigations, particularly when the quantity of interest is biomass. By examining 20 antimicrobials spanning 11 classes, the work shows that discrepancies between CFU and luminescence are often biological (filamentation, Viable But Not Culturable). The work provides a convincing view of how these three common measurements (luminescence, optical density, and CFU) relate to one another across a range of drug treatments, although testing on clinical isolates could be of further benefit.

    2. Reviewer #2 (Public review):

      Summary:

      In antibiotic research, accurately measuring decreases in bacterial populations is essential. The authors conducted a comprehensive evaluation of the luminescence assay, a commonly used but previously under-quantified method, benchmarking it against the gold-standard CFU counting approach. They found that luminescence measurements generally aligned with CFU results but sometimes reported slower decline rates for certain antimicrobials. These discrepancies were linked to differences in how the two methods capture biomass and colony formation, which vary with the antimicrobial's mechanism of action. The study demonstrates that luminescence assays can serve as a high-throughput alternative to labor-intensive CFU counting, provided their limitations are understood and corrected.

      Strengths:

      The authors developed a mathematical model to partially correct luminescence-based measurements, making the approach broadly applicable to several commonly used antibiotics. They also analyzed antibiotic-treated single-cell morphologies and linked filamentation to bulk luminescence signals. This analysis helped define the range of drug conditions under which luminescence assays provide reliable estimates of bacterial dynamics.

      They extensively evaluated the method using 20 antibiotics and one antimicrobial peptide, encompassing many of the most commonly used agents and experimental factors (e.g. treatment time) typically considered in antibiotic research.

      Comments on revised version:

      No further comments. The authors have adequately addressed my concerns.

    3. Reviewer #3 (Public review):

      Summary:

      This preprint proposes luxCDABE-based luminescence as a high-throughput alternative (or complement) to CFU time-kill assays for estimating antimicrobial rates of population change at super-MIC concentrations, by comparing luminescence- and CFU-derived rates across 20 antimicrobials (22 assays) and attributing divergences primarily to filamentation (luminescence closer to biomass/volume than cell number) and changes in culturability / carryover (CFU undercounting viable cells).

      Strengths:

      The authors do not merely report discrepancies; they experimentally validate the biological causes. Specifically, they successfully attribute the slower decline of luminescence in certain drugs to bacterial filamentation (maintaining biomass despite halted division) and the rapid decline of CFU in others to loss of culturability or carryover effects.

      The inclusion of 20 antimicrobials spanning 11 classes provides a robust dataset that allows for broad categorization of drug-specific assay behaviors.

      The study critically exposes flaws in the "gold standard" CFU method, specifically regarding antimicrobial carryover (demonstrated with pexiganan) and the potential for CFU to overestimate cell death in the presence of VBNC (viable but non-culturable) states induced by drugs like ciprofloxacin.

      The use of chromosomal integration for the lux operon to minimize plasmid copy-number effects and the validation of linearity between light intensity and cell density establish a solid technical foundation.

      In summary:<br /> Muetter et al. provide a compelling argument that luminescence is a reliable, high-throughput alternative to CFU for super-MIC investigations, particularly when the quantity of interest is biomass. The paper effectively warns researchers that discrepancies between CFU and luminescence are often biological (filamentation, VBNC) rather than methodological failures.

      Comments on revised version:

      The revised version addressed my comments well.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study examines how luminescence can be used to measure bacterial population dynamics during antimicrobial treatment by comparing it directly with optical density and colony counts. The authors aim to determine when luminescence reflects changes in population size and when it instead captures metabolic or physiological states induced by drug exposure. By generating parallel datasets under controlled conditions, the work provides a detailed view of how these three common measurements relate to one another across a range of drug treatments.

      Strengths

      The study is technically strong and thoughtfully designed. Measuring luminescence, optical density, and colony counts from the same cultures allows the authors to make clear and informative comparisons between methods. The data are compelling, and the analyses highlight both agreements and divergences in a way that is easy to interpret. The manuscript also succeeds in showing why these divergences arise. For example, the observation that filamentation and metabolic shifts can sustain luminescence even when colony counts drop provides valuable information on how different readouts capture distinct aspects of bacterial physiology. The writing is clear, the figures are effective, and the work will be useful for researchers who need high-throughput approaches to quantify microbial population dynamics experimentally.

      Weaknesses:

      The study also exposes some inherent limitations of luminescence-based measurements. Because luminescence depends on metabolic activity, it can remain high when cells are damaged or unable to resume growth, and it can fall quickly when drugs disrupt energy production, even if cells remain physically intact. These properties complicate interpretation in conditions that induce strong stress re-sponses or heterogeneous survival states.

      In addition, the use of drug-free plates for colony counts may overestimate survival when filamented or stressed cells recover once the antibiotic is removed, making differences between luminescence and colony counts harder to attribute to killing alone. Finally, while the authors discuss luminescence in the context of clinically relevant concentration ranges, the current implementation relies on engineered laboratory strains and does not directly demonstrate applicability to clinical isolates. These limitations do not detract from the technical value of the work but should be kept in mind by readers who wish to apply the method more broadly.

      We thank the reviewer for reading our paper thoroughly and for the helpful feedback.

      Luminescence limitations. We agree that the lack of a direct link between light intensity and a population property such as biomass or cell number is the main limitation of the luminescence method. To further emphasise this, we have expanded the Discussion in the revised manuscript.

      Drug-free plates. The use of drug-free plates is intentional. As we measure a time series, the question at each point is how many cells are alive at each time point. Cells that are stressed but viable at time t contribute correctly to the count at t. How long they survive under the respective treatment is captured by the subsequent timepoints.

      Filaments. Recovery of plated filamented cells should not inflate this estimate. A single plated filamentous cell is expected to yield either zero (death before division) or one single colony, regardless of in how many parts it separates, as all descendants are part of the same cluster. However, if the cells divide before plating, CFU can overestimate survival. Having that said, we have no indication that this occurred in our experiments, since in all observed discrepancies, CFU-based estimates were equal to or lower than those obtained from luminescence and the time cells spent in dilution was kept short.

      Clinical applicability. We agree with the reviewer that the method is not practical for ad-hoc pharmacodynamic studies of clinical isolates. What we instead provide is an E. coli-based model system to explore clinically relevant treatment conditions, which we address in the revised manuscript. We believe that constructing analogous bioluminescent model strains in other clinically relevant species would be a valuable direction for future work.

      Reviewer #2 (Public review):

      Summary:

      This preprint proposes luxCDABE-based luminescence as a high-throughput alternative (or complement) to CFU time-kill assays for estimating antimicrobial rates of population change at super-MIC concentrations, by comparing luminescence- and CFU-derived rates across 20 antimicrobials (22 assays) and attributing divergences primarily to filamentation (luminescence closer to biomass/volume than cell number) and changes in culturability/carryover (CFU undercounting viable cells).

      Strengths:

      The authors do not merely report discrepancies; they experimentally validate the biological causes. Specifically, they successfully attribute the slower decline of luminescence in certain drugs to bacterial filamentation (maintaining biomass despite halted division) and the rapid decline of CFU in others to loss of culturability or carryover effects.

      The inclusion of 20 antimicrobials spanning 11 classes provides a robust dataset that allows for broad categorisation of drug-specific assay behaviours.

      The study critically exposes flaws in the “gold standard” CFU method, specifically regarding antimicrobial carryover (demonstrated with pexiganan) and the potential for CFU to overestimate cell death in the presence of VBNC (viable but non-culturable) states induced by drugs like ciprofloxacin.

      The use of chromosomal integration for the lux operon to minimise plasmid copy-number effects and the validation of linearity between light intensity and cell density establish a solid technical foundation.

      Weaknesses:

      The study is conducted exclusively using Escherichia coli. While E. coli is a standard model organism, the paper claims to evaluate luminescence as a generalisable high-throughput tool. Many of the discrepancies observed are driven by filamentation. However, distinct morphological responses occur in other critical pathogens (e.g., Staphylococcus aureus does not filament in the same way).

      The authors propose that luminescence data can be corrected using microscopyderived volume data to better align with CFU counts. The primary appeal of luminescence is high-throughput efficiency. If a researcher must perform timelapse microscopy to calculate cell volume changes to “correct” their luminescence data, the high-throughput advantage is lost.

      The paper argues that for ciprofloxacin, CFU underestimates viability because cells remain intact and impermeable to propidium iodide. While the cells are metabolically active and membrane-intact, if they cannot divide to form a colony (even after drug removal/dilution), their clinical relevance as “living” pathogens is debatable.

      Some other comments:

      The use of a population dynamical model to simulate filamentation effects is excellent. The finding that light intensity tracks volume ($\psi_V$) better than cell number ($\psi_B$) is a key theoretical contribution.

      The model assumes linear elongation. The authors should briefly comment on whether this holds true for the specific drug mechanisms tested (e.g., PBP inhibition vs. DNA gyrase inhibition).

      The use of bootstrapping to estimate rate distributions is appropriate and robust.

      Conclusion:

      Muetter et al. provide a compelling argument that luminescence is a reliable, highthroughput alternative to CFU for super-MIC investigations, particularly when the quantity of interest is biomass. The paper effectively warns researchers that discrepancies between CFU and luminescence are often biological (filamentation, VBNC) rather than methodological failures.

      We thank the reviewer for reading our paper thoroughly and for the helpful feedback.

      Generalisability. We agree that the alignments and divergences reported for specific drugs may not transfer directly to other species, which may elongate differently (e.g. cocci) or show different physiological responses to treatment. Constructing analogous model strains — for example based on S. aureus to cover a broader range of morphologies and clinically relevant species would therefore be an interesting follow-up project, and we have adjusted the Discussion to make this clearer. We nevertheless believe that the broader conclusions (larger cells emit more light) of the paper likely hold across species.

      Volume correction. We agree that requiring microscopy would undermine the high-throughput advantage of the luminescence assay. It was not our intention to propose this as a practical approach, nor to imply that the luminescence signal needs a correction. Taken on its own, the signal can be interpreted as the cumulative metabolic output of the population, which is closely linked to biomass, and that measure is valuable in itself for many applications. We used the volume correction only to demonstrate that luminescence tracks biomass more closely than cell number: by adjusting the luminescence distribution with the measured volume change, it moves towards the CFU distribution. We have revised the Discussion to prevent this from being misunderstood as a required step.

      Culturability vs. clinical relevance. We agree that the dynamics of culturable cells are highly relevant, especially in a clinical context. Our aim was to explain the observed differences between CFU and luminescence by highlighting that culturability and viability are not always identical, without implying that one measure is inherently superior to the other — we leave it to the reader to decide which metric best suits their needs.

      Linear elongation. The model assumes linear elongation for mathematical convenience, which, depending on the specific strain and drug mechanism, could be incorrect. Its purpose is to demonstrate that a shift of the mean cell volume to a new, higher equilibrium under treatment can cause an initial peak in the luminescence signal despite a declining population. This remains true for non-linear elongation models, though the shape, height and position of the peak may change. We have adjusted the Results to make this clearer.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors present luminescence as a practical measurement of population decline under antibiotic exposure. One aspect that could be clarified is how the method behaves when tolerance arises from phenotypic heterogeneity, such as the presence of small, metabolically quiet survivors. Because luminescence reflects metabolic activity and biomass, the signal will be dominated by metabolically active cells, making rare tolerant subpopulations difficult to detect. A short discussion of how luminescence performs in these heterogeneous scenarios, and whether complementary assays are needed to capture long-lived tolerant cells, would strengthen the manuscript.

      Yes, that is a valid concern and we thank the reviewer for raising this point.

      Heterogeneity in cell-specific luminosity alone does not bias population-level rate estimates. A bias can arise, however, when specific luminosity correlates with a second factor — most importantly, the decline rate under treatment.

      We agree with the reviewer’s suggestion that brighter cells plausibly die faster than tolerant, metabolically quiet ones. When one subpopulation dominates the light signal, we expect minimal bias, as the rate estimate primarily reflects that subpopulation. However, in a transition phase when both subpopulations contribute roughly equally to the light signal, luminescence likely overestimates the decline.

      We added a corresponding caveat to the Discussion (lines 581–583).

      (2) The manuscript shows that filamentation can influence ψ_I by altering biomass and metabolic activity independently of cell number. However, antibiotic exposure can also trigger other stress responses and metabolic shifts that change energy fluxes, redox balance, and biosynthetic activity. Since luminescence depends on metabolic state and substrate availability, these additional physiological transitions may also affect ψ_I in ways not directly tied to birth or death processes. It would be useful to comment on whether such responses, beyond filamentation, are likely to influence luminescence dynamics across different drug classes or treatment conditions.

      We thank the reviewer for raising this point and agree that there is no biological law strictly linking luminosity to a single population property such as biomass or cell number, and changes in the metabolism most likely affect Ψ<sub>I</sub> as well.

      Transitioning to a new metabolic steady state biases Ψ<sub>I</sub>; once the new steady state is reached, however, the rate estimate should no longer be affected.

      Looking across drug classes, drugs that primarily lyse cells (polymyxins and, to a lesser degree, beta-lactams targeting PBP1) did not show noticeable deviations between Ψ<sub>I</sub> and Ψ<sub>CFU</sub>, and — perhaps counterintuitively — neither did ribosome-inhibiting drugs.

      For the remaining cases, we were able to attribute part of the discrepancy between Ψ<sub>I</sub> and Ψ<sub>CFU</sub> to changes in biomass or loss of culturability, though drug-induced metabolic changes may also contribute to the residual differences.

      We clarify this in the Discussion (lines 569–579).

      (3) The authors quantify survival using colony counts on drug-free medium. Because filamentation can be a reversible state that persists during antibiotic exposure, plating on drug-free medium may capture recovery potential rather than in-treatment viability. Filamented or stressed cells that cannot divide in the presence of a drug may nevertheless form colonies once the drug is removed. Clarifying how this recovery step affects ψ_CFU would help readers interpret differences between luminescence-based and colony-based measurements, especially in cases where transient tolerant states are present.

      We thank the reviewer for raising this point.

      Our CFU assay estimates the number of culturable cells at each time point; the rate Ψ<sub>CFU</sub> is then inferred from how this number changes across time points. Plating on drug-free medium is intentional, as it maximises the probability that a culturable cell is detected at each snapshot. Whether those cells would have continued dividing or died under continued treatment is captured by the subsequent time points.

      Filamentation interacts with the probability of colony formation in several, partly opposing ways:

      (1) It can increase the death rate, as for ceftazidime and cefepime, which is part of the kill effect captured by Ψ<sub>CFU</sub>;

      (2) Entanglement between filaments may reduce the number of colonies per plated bacterium;

      (3) Conversely, if a filament divides upon drug removal, its fragments form a cluster that — stochastically — is very likely to produce one (but not multiple) colony.

      The only scenario in which CFU could overestimate bacterial density is if a filament separates into individual cells in the liquid phase before plating; we have no indication that this occurred in our experiments.

      We addressed this concern in our response to the public comment.

      (4) A brief discussion comparing luminescence to fluorescent reporter systems could be helpful. Fluorescent proteins typically require a chromophore maturation step before becoming detectable, which introduces a delay between the underlying cellular event and the appearance of the signal. In contrast, as far as I understand, lux reporters emit light immediately once the enzymatic components and substrates are present, without a maturation stage. Highlighting this distinction may help readers understand why luminescence is well-suited for tracking rapid changes in population physiology under antibiotic exposure. However, the manuscript also notes that luminescence can lag slightly behind very rapid killing (particularly for AMPs), but the temporal dynamics of signal shutdown are not explored in detail. Because lux reflects metabolic activity rather than viability, a short delay between irreversible damage and the loss of light is biologically expected. It may help readers if the authors could expand on the mechanism underlying this delay in order to clarify when ψ_I is likely to track true biomass decline and when residual metabolic activity might mask early killing events.

      On fluorescent reporters:

      We thank the reviewer for this suggestion.

      Under some conditions, change rates can also be measured using fluorescence, provided the number of fluorescent molecules per bacterium remains constant. This requires a balance between production, maturation, degradation and dilution, which is only established if the growth rate and conditions remain constant over a sufficiently long period (typically hours).

      For measuring population decline, however, the key issue is that cell death does not inactivate fluorescent proteins: once matured, they emit independently of the cell’s metabolic state and decay only with the protein’s half-life, which is typically slower than the kill rates of interest.

      We added a clarification to the Introduction (lines 58–60).

      On the lux signal lag:

      We thank the reviewer for raising this point. The short lag between luminescence and CFU decline could in principle arise from two mechanisms: (i) luminescence declining more slowly than the actual cell number (residual light from dead cells), or (ii) CFU declining more steeply than the actual cell number (damaged but still viable cells failing to form colonies).

      Mechanism (i) splits into two sub-cases:

      (i.a) Dead but impermeable — the lux reaction could in principle continue for a short while if enough components are retained in the cell. However, a metabolically active, impermeable cell is difficult to classify as dead in the first place, making this scenario conceptually awkward.

      (i.b) Dead and permeable (lysed) — the lux components dilute into the medium, and by mass-action the reaction rate should drop rapidly (though not instantly). Any residual signal after lysis should therefore be short-lived.

      Mechanism (ii) — damaged (e.g. permeable) cells may be particularly sensitive to plating on agar (e.g. due to oxidative stress), resulting in a declining probability of colony formation.

      In our case, the discrepancy was observed specifically for pexiganan, where cells can be assumed to lyse, making (i.b) and/or (ii) the likely explanations. Based on our experimental data, we cannot distinguish between these possibilities and therefore limit ourselves to reporting the observed discrepancy.

      We have moved the interpretation from the Results to the Discussion (lines 523–542) and expanded the discussion there.

      (5) In lines 85–89, the authors state that “high-throughput OD and luminescence measurements at sub-MIC concentrations provide valuable insights into drug effects on growth rates, [but] the super-MIC range is clinically more relevant,” and they present luminescence as a way to investigate super-MIC population dynamics. While super-MIC behaviour is indeed important for pharmacodynamics and resistance evolution, it is not clear that the specific luminescence implementation used here has direct clinical relevance. The study relies on a chromosomally integrated reporter in a laboratory strain, and the manuscript does not demonstrate that this approach can be applied to clinical isolates or diagnostic workflows. It may be helpful to moderate the claim of “clinical relevance” and frame the method more clearly as a high-throughput experimental tool that can inform clinically relevant questions, rather than as an assay ready for clinical application.

      We agree and have moderated the framing accordingly (lines 100–104).

      Reviewer #2 (Recommendations for the authors):

      (1) The conclusions regarding “biomass vs. cell number” may not apply equally to non-rod-shaped bacteria or species with different stress responses. The authors must explicitly discuss this limitation in the Discussion.

      The broad conclusion that bigger cells emit more light likely holds across morphologies, since it rests on the principle that more cellular material means more metabolic activity and therefore more light. The quantitative relationship between cell size and luminosity, however, may differ across species, shapes and conditions, for two reasons. First, chromosome copy number: whether drug-induced morphological changes are accompanied by chromosome replication and therefore an increase in lux operon copy number — varies across species and drug mechanisms. Second, the surface-to-volume ratio likely modulates mass-specific metabolism; some morphological changes preserve it (e.g. purely lateral elongation) while others do not.

      The more specific conclusions about which drug classes produce alignment or divergence between CFU and luminescence may also not transfer directly, as drug mechanisms can act differently across species.

      We already note this limitation in the Discussion (lines 594– 597) and have expanded the wording there.

      (2) The manuscript should clarify that luminescence is a superior metric for biomass without correction, rather than framing the volume correction as a necessary step to mimic CFU. The divergence should be embraced as a feature (biomass tracking), not a bug that needs fixing via labor-intensive microscopy.

      We agree with the framing and will make it clearer; it was actually our intention to clarify which method does what, rather than judge one as better or worse.

      We removed the “correction” sentence from the Discussion to make this clearer.

      (3) The authors should refrain from definitively stating CFU “underestimates” viability and instead use more precise terminology, such as “reproductive capability” vs. “metabolic integrity.”

      We agree with the reviewer that measuring culturability is a property, not a flaw, of CFU. Our intention was to emphasise that when CFU is used as a proxy for viability (which it often is), it can yield lower values than the actual number of survivors. We tried to make that distinction explicit in the manuscript (e.g. in lines 317–322).

      We would also like to note that in the case of antimicrobial carryover, CFU can genuinely underestimate culturability itself, not only viability.

      Regarding the suggested reproductive capability vs. metabolic integrity framing: we agree that metabolism and luminescence are closely linked. What held us back from drawing that link directly is that metabolism is hard to quantify, being the cumulative output of a diverse set of processes.

      (4) The model assumes linear elongation. The authors should briefly comment on whether this holds true for the specific drug mechanisms tested (e.g., PBP inhibition vs. DNA gyrase inhibition).

      Linear elongation is a mathematically convenient simplification whose only purpose in the model is to allow the population to converge to a new equilibrium volume under treatment. Assuming constant volume-specific luminosity, we showed that this produces an initial peak in light intensity before the signal declines in parallel with Ψ<sub>B</sub>. The exact shape, height and position of this peak depend on the volume growth model used, but the qualitative pattern — peak followed by parallel decline — holds for other growth models as well. We now clarify this in lines 230–235.

      (5) The authors suggest the carryover effect is due to a delay between cell death and cessation of luminescence. This “lag time” is a critical physical constraint of the lux system (likely related to ATP depletion or enzyme decay) and should be quantified or discussed in more detail as a fundamental “speed limit” for the assay.

      The origin of the lag between luminescence and CFU is an interesting question, but one we cannot definitively answer. We can, however, discuss the potential mechanisms:

      A dead but impermeable cell could in principle continue to emit residual light for some time. We note, though, that calling a metabolically active, impermeable cell “dead” is a question of definition we would rather not discuss here.

      In our case, the discrepancy was observed specifically for pexiganan, where cells can be assumed to lyse. Under lysis, the lux components dilute quickly into the medium, and by mass-action the reaction rate should drop rapidly — though not necessarily instantaneously.

      A plausible alternative to a delayed cessation of the light signal is that the probability of colony formation drops rapidly after permeabilisation, for example because permeable cells are sensitive to oxidative stress when plated on agar.

      Based on our experimental data we cannot distinguish between these mechanisms, so we limit ourselves to reporting the observed discrepancy. We have moved the interpretation from the Results to the Discussion (lines 523–542) and expanded on the candidate mechanisms there.

      Additional revisions

      Beyond the changes prompted by the reviewers’ comments, we made the following revisions to the supplementary information:

      We corrected the Λ matrix (converted row 2, col 4 from 0 → 2)

      We removed the line numbering

    1. eLife Assessment

      This valuable study introduces MULTI i<sup>2</sup>, a robust and high-throughput method to measure Plasmodium falciparum viability in the presence of drugs. This new assay offers significant time savings over the traditional Parasite Reduction Rate (PRR) assay and should enable faster screening of drug combinations, which is urgently needed in the field. The assay is well validated, with convincing data showing it can reproduce known drug interactions and identify new interaction patterns.

    2. Reviewer #1 (Public review):

      Summary:

      The authors describe a clever genetic system based on rapamycin-inducible expression of a beta-galactose reporter. The authors compare this spectrophotometer-based readout to the parasite reduction rate version 2 (PRRv2) recently described by some of the same authors and based on incorporation of [3H]-hypoxanthine. The results are generally comparable, with some differences for slower-acting compounds. The authors report that this format is better suited for higher-throughput studies and requires less time to quantify the time-dependent onset of parasiticidal action compared with the PRRv2.

      Strengths:

      This is a very well-executed and well-described body of work with a comprehensive set of analyses.

      Weaknesses:

      The authors should revise their text to also describe other methods used to quantify parasite growth. This method saves time compared to the PRRv2 but is too complex for simple screening of antiplasmodial activity of agents tested alone. Its value lies in assessing the speed of action of compounds tested in combination.

      There are a number of areas for improvement:

      (1) Many antimalarials have quite specific times of action. Are these MULTI-i2 assays, and the comparator PRRv2 assays, conducted with asynchronous cultures? This should be described in the methods and referred to in the text (apologies if I missed some references).

      (2) The authors correctly state that flow cytometry-based readouts, such as with MitoTracker alone, can limit throughput and that MitoTracker alone can produce spurious results. The authors should cite work from other labs that combine MitoTracker with a nuclear dye, such as SYBR Green I. I think others have also been used, such as YoYo-1, which overcomes the limitations of using MitoTracker alone. Also, many labs use a nuclear dye such as SYBR Green I in a spectrophotometer-based format that enables rapid processing of plates at scale (96, 384, or even 1536 wells per plate). Luciferase-based screens have also been used in large-scale screening campaigns. The introduction should cite these various approaches, especially as the MULTI-i2 method is quite a complex screen with an initial period of drug exposure (up to 3 days) followed by a five-day phase initiated by rapamycin addition to induce expression of the beta-gal sensor.

      (3) It would be helpful for authors to provide some indication of the cost comparison between the PPRv2 and MULTI-i2.

      (4) Also, the authors should indicate whether these reagents will be deposited in a repository such as BEI Resources. They should also indicate conditions for other groups to request these materials, such as whether an MTA is required.

      (5) The pharmacological models are interesting, but likely well out of the range of expertise of many labs. Has code been deposited into public repositories that make it possible for other labs to implement these analyses?

    3. Reviewer #2 (Public review):

      Summary

      Antimalarial combination therapy is the standard of care for malaria, a disease that impacts hundreds of millions of people annually. Combination therapy is crucial for effectively treating the disease and delaying the emergence of drug resistance. Despite the importance of choosing appropriate partner antimalarials for combination therapy, drug interactions are typically evaluated late in the course of drug development. Standard in vitro assays that determine synergistic, antagonistic, or additive interactions between drug combinations rely on measuring inhibition of parasite proliferation, which is inadequate for translation to pharmacodynamic models for parasite clearance in the patient. Direct measurement of parasite viability under drug treatment has previously relied on methods that are labor and resource-intensive, limiting applications to single compounds and single concentrations. Here, Hellingman et al make use of an inducible chemiluminescence reporter to measure cell viability and apply this novel approach to quantify drug interactions. The methodology is a significant improvement upon prior methods, requiring significantly fewer resources, half the time, and substantially less handling than the standard PRRv2 assay, whilst maintaining high resolution and sensitivity.

      They assess the limit of detection for the improved method and cross-reference their results for single drugs at a single concentration with the currently standard PRRv2 assay. The authors next established analytical methods to characterize the impact of drug combinations on parasite viability using the GDPI pharmacodynamic model and compared their MULT-i2 assay to the prior cPRR approach. Their refined workflow allowed them to comprehensively evaluate the known synergistic combination between atovaquone and proguanil with greater resolution than the comparable cPRR assay and identified additional interaction parameters between the fast-acting antimalarials piperaquine and pyrimethamine. Overall, the authors demonstrate that their inducible lacZ system provides significant advantages compared with prior approaches to determine parasite viability. They convincingly demonstrate the strengths of their approach by characterizing two antimalarial combinations at much greater resolution than previously possible with prior methods. The system and methods established here will be particularly useful for evaluating novel antimalarial combinations with chemical series in preclinical evaluation and to optimize future antimalarial therapies.

      Strengths:

      The streamlined approach relies on induction of the lacZ enzyme only after drug washout. As opposed to when stably expressed, this allows the authors to estimate parasite viability without undergoing serial dilutions to estimate viable parasite titers. This innovation vastly reduced resource and time intensity, enabling greater throughput for parasite viability estimation. The established methodology and analysis pipeline enabled the testing of 49 drug combinations for parasite viability in the MULT-i2 assay compared to only 9 in the conventional cPRR assay. This provided improved resolution in the ability to estimate drug combination parameters in a pharmacodynamic model. The ability to comprehensively characterize combination pharmacodynamic properties in vitro will have important implications for downstream modelling of in vivo combinations, and for optimizing future antimalarial combination therapies.

      The authors made good use of modelling and AICc for parametric estimation and model evaluation to demonstrate the advantages of the richer dataset afforded by the MULT-i2 assay.

      Weaknesses:

      The authors correctly identified a range of confounding effects that lead to artefacts in their assay results when compared to the cPRR assay. For instance, the authors observed reduced signal at high parasite density during recovery due to overgrowth and likely enzyme degradation, and suggested residual signal may remain from non-proliferating sexual stage parasites surviving drug treatment that would not be detected in the cPRR assay.

      Measurement of parasite viability in the MULT-i2 assay was achieved by extrapolating the chemoluminescence signal to that of a serial dilution of parasites made at the initiation of drug treatment. How did the authors account for differing levels of enzyme expression at early (e.g., ring) vs late stage parasites (trophozoite or schizonts)? Were cultures synchronized prior to initiation of assays? Could differences in life-cycle progression following drug treatment be an additional confounding factor that may account for differences with the PRRv2 assay?

      The addition of an inducible element is an improvement of their earlier lacZ/β-galSENSOR (PMID: 41575867); however, the authors fail to explain why this is an improvement and how this adds additional merit over the initial system. While the authors compare their new assay to the PRRv2, they fail to compare it to their own non-inducible lacZ/β-galSENSOR system. Their non-inducible system already showed superiority to the cPRR assays, and it would be good to show how they compare and what the advantages of the new system are over the old. e.g., how is the signal-to-noise improved? How does the sensitivity compare? How quickly does the can the signal be detected after induction? They show signal after 48h, but it would be very useful to the community to look at earlier timepoints as well and compare them to the uninduced line and a line that has been induced 48h earlier to match the expression patterns throughout the lifecycle (something like 2h,4h,6h, 12h, and 24h).

      Is the chemiluminescence signal for the i-lacZ induced parasites comparable to the stably expressed lacZ parasites previously characterized by the group? If so, do the authors consider this inducible iteration a complete replacement for PRR assays?

    4. Reviewer #3 (Public review):

      In this manuscript, the authors strived to develop a highly efficient drug survival assay for in vitro cultured human malaria parasites P. falciparum. This was done by generating a transgenic P. falciparum line using a creLox strategy that allows detection of (presumably) viable parasites by a β-lactamase assay. To estimate the Limit of quantification of the recombined P. falciparum NF54i-lacZ, the authors ultimately designed a protocol in which viable parasites are detected by the luminescence of β-D-galactoside generated by β-lactamase within the transgenic parasites. For this, the parasite must be incubated with rapamycin for 120 hours to induce CreLox recombinase, which places β-lactamase under an active promoter. Using this assay, termed MULTI-i2, the author shows interactions between two antimalarial drug pairs that were previously demonstrated by another assay. In the case of pyronaridine and piperaquine pair, the NULT-i2 assay generated some additional insights compared to the previous assay, presumably by virtue of including more concentration datapoints. In conclusion, the authors argue that the MULTI-i2 assay is much less resource-intensive and time-consuming and can be applied on a large scale at a much lower cost and with the highest efficiency.

      Overall, the data generated in this manuscript are clear and well represented, and I am convinced that MULTI-i2 provides yet another of many drug assays for malaria parasites and could be put to good use. However, I struggle to fully appreciate the merit of his study, as the manuscript reads more like a technical document than a scientific study.

      I particularly lack an understanding of the strengths and weaknesses/limitations of the MULTI-i2 methodology and, thus, its applicability. I also do not fully appreciate the need for such an elaborate luminescence-based experimental setup. It would be good if some of these issues were addressed.

      Specifically:

      (1) The whole assay is based on detecting parasites by luminescence after 120 hr (5 days) after drug exposure. During that time, presumably the parasites that survived the drug pressure regrow to a detectable level and, at the same time, perform efficacious CreLox-based recombination to produce β-D-galactoside for detection. Is this necessary? How superior is this detection method to other methods, such as Fluorescence-assisted Cell Sorting (FACS), etc? Moreover, the 5-day growth-CreLox-β-D-galactoside production could introduce a series of confounding effects. In my view, more studies (beyond comparisons with a single existing method) would be useful for understanding this entire process.

      (2) Related to that above, how would MULTI-i2 perform in case of drugs that do not necessarily kill all parasites, such as artemisinin? In the case of artemisinin, it is becoming evident that at least a small fraction of the parasite revives after treatment via a temporary dormancy state. This has, in fact, also been shown for other drugs such as mefloquine, pyrimethamine, etc. Would such a situation produce a range of false readings? In general, in its current state, it is hard to see what the limitations of this method are, which makes it hard to decide whether to use it for a particular application.

      (3) Given the stated cost and labor efficiency of MULTI-i2, it is disappointing to see only two applications for two drug pairs: atovaquone/proguanil and piperquine/pyronaridine, for both of which their interactions were already known. The manuscript would benefit greatly if the authors demonstrated more drug interactions and identified (and ultimately validated) new ones. This would certainly make MULT-i2 method more attractive. In particular, it would be nice to see if one could use MULTI-i2 for studies of triple combinations as enthusiastically suggested.

      (4) Throughout the manuscript, the authors claim that MULTI-i2 is considerably less expensive and can be done much faster than previous methods. In my view, this is not exactly a scientific argument. The cost of an assay depends heavily on the cost of reagents and labor, which are subject to market price fluctuations. The efficiency and time consumption can very much depend on laboratory organization, etc. Unless the author could specifically demonstrate where and how these assays are cheaper and faster, I suggest not discussing this.

    5. Author response:

      Reviewer 1 (Public review):

      Summary:

      The authors describe a clever genetic system based on rapamycin-inducible expression of a beta-galactose reporter. The authors compare this spectrophotometer-based readout to the parasite reduction rate version 2 (PRR v2) recently described by some of the same authors and based on incorporation of [3H]-hypoxanthine. The results are generally comparable, with some differences for slower-acting compounds. The authors report that this format is better suited for higher-throughput studies and requires less time to quantify the time-dependent onset of parasiticidal action compared with the PRR v2.

      Strengths:

      This is a very well-executed and well-described body of work with a comprehensive set of analyses.

      Weaknesses:

      The authors should revise their text to also describe other methods used to quantify parasite growth. This method saves time compared to the PRR v2 but is too complex for simple screening of antiplasmodial activity of agents tested alone. Its value lies in assessing the speed of action of compounds tested in combination.

      We thank reviewer 1 for the supportive feedback and for raising some important points.

      Many antimalarials have quite specific times of action. Are these MULT-i<sup>2</sup> assays, and the comparator PRR v2 assays, conducted with asynchronous cultures? This should be described in the methods and referred to in the text (apologies if I missed some references).

      We thank the reviewer for this important comment. Both, the MULT-i<sup>2</sup> and PRR v2 assays were performed using asynchronous parasite cultures. This information is included in the Methods section together with the relevant references. To improve clarity, we will also explicitly state this in the main text.

      The authors correctly state that flow cytometry-based readouts, such as with MitoTracker alone, can limit throughput and that MitoTracker alone can produce spurious results. The authors should cite work from other labs that combine MitoTracker with a nuclear dye, such as SYBR Green I. I think others have also been used, such as YoYo-1, which overcomes the limitations of using MitoTracker alone. Also, many labs use a nuclear dye such as SYBR Green I in a spectrophotometer-based format that enables rapid processing of plates at scale (96, 384, or even 1536 wells per plate). Luciferase-based screens have also been used in large-scale screening campaigns. The introduction should cite these various approaches, especially as the MULT-i<sup>2</sup> method is quite a complex screen with an initial period of drug exposure (up to 3 days) followed by a five-day phase initiated by rapamycin addition to induce expression of the beta-gal sensor.

      We thank the reviewer for this helpful suggestion. In the Introduction we will mention and describe alternative approaches for assessing parasite viability. This will also include the work by Maiga et al., which combines MitoTracker with a nuclear dye to improve the reliability of flow cytometry-based readouts. We will revise the text to explicitly mention the use of dual straining to make this discussion more explicit.

      We agree that several additional methods, such as luciferase-based reporter systems, have been successfully applied in antimalarial screening. However, these approaches are primarily designed to assess parasite growth inhibition rather than directly measuring parasite viability after drug exposure, which is the focus of the present study. Readout methods used to assess parasite viability in a PRR assay setup are so far based on HRP2-ELISA (de Carvalho et al.), MitoTracker and SYBR green staining (Maiga et al.) and [<sup>3</sup>H]-hypoxanthine incorporation (Sanz et al.; Walz et al.) as cited in the manuscript. Many other readout methods to assess parasite growth have other limitations as briefly discussed in Hellingman et al, 2024. A comprehensive comparison and review of all available readout methods would therefore be beyond the scope of this manuscript.

      It would be helpful for authors to provide some indication of the cost comparison between the PPR v2 and MULT-i<sup>2</sup>.

      We thank the reviewer for this valuable suggestion. We agree that a comparison of the costs associated with the PRR v2 and MULT-i<sup>2</sup> assays would be informative, but while the consumable costs provide one measure of assay expense, we consider the reduction in hands-on time and the simplified workflow to be the main contributors to the overall cost advantage of the MULT-i<sup>2</sup> assay. These reductions in labor requirements are subject to large regional differences and impossible for us to access. Nevertheless, together with the increased throughput and the reduced labor, make the MULT-i<sup>2</sup> assay more cost-effective for larger-scale applications compared with the PRR v2 assay.

      Also, the authors should indicate whether these reagents will be deposited in a repository such as BEI Resources. They should also indicate conditions for other groups to request these materials, such as whether an MTA is required.

      We thank the reviewer for this important suggestion. The engineered parasite line will be made available for non commercial use to other researchers upon request. An MTA will be required excluding commercial use of the provided strains. The detailed code used for data analysis is available upon request, and an example code file has already been included as a Supplementary File.

      The pharmacological models are interesting, but likely well out of the range of expertise of many labs. Has code been deposited into public repositories that make it possible for other labs to implement these analyses?

      We thank the reviewer for this valuable comment. We agree that implementation of pharmacological modeling approaches can represent a barrier for laboratories without prior experience in pharmacometric analysis, particularly due to the requirement for specialized software such as NONMEM. To facilitate implementation, an example code is provided in the Supplementary File. The final model was developed using a forward–backward selection approach for parameter estimation and model refinement as described in the Methods section. These additions should help other researchers adapt the approach to their own datasets.

      Reviewer 2 (Public review):

      Summary

      Antimalarial combination therapy is the standard of care for malaria, a disease that impacts hundreds of millions of people annually. Combination therapy is crucial for effectively treating the disease and delaying the emergence of drug resistance. Despite the importance of choosing appropriate partner antimalarials for combination therapy, drug interactions are typically evaluated late in the course of drug development. Standard in vitro assays that determine synergistic, antagonistic, or additive interactions between drug combinations rely on measuring inhibition of parasite proliferation, which is inadequate for translation to pharmacodynamic models for parasite clearance in the patient. Direct measurement of parasite viability under drug treatment has previously relied on methods that are labor and resource-intensive, limiting applications to single compounds and single concentrations. Here, Hellingman et al make use of an inducible chemiluminescence reporter to measure cell viability and apply this novel approach to quantify drug interactions. The methodology is a significant improvement upon prior methods, requiring significantly fewer resources, half the time, and substantially less handling than the standard PRR v2 assay, whilst maintaining high resolution and sensitivity.

      They assess the limit of detection for the improved method and cross-reference their results for single drugs at a single concentration with the currently standard PRRv2 assay. The authors next established analytical methods to characterize the impact of drug combinations on parasite viability using the GDPI pharmacodynamic model and compared their MULT-i<sup>2</sup> assay to the prior cPRR approach. Their refined workflow allowed them to comprehensively evaluate the known synergistic combination between atovaquone and proguanil with greater resolution than the comparable cPRR assay and identified additional interaction parameters between the fast-acting antimalarials piperaquine and pyrimethamine. Overall, the authors demonstrate that their inducible lacZ system provides significant advantages compared with prior approaches to determine parasite viability. They convincingly demonstrate the strengths of their approach by characterizing two antimalarial combinations at much greater resolution than previously possible with prior methods. The system and methods established here will be particularly useful for evaluating novel antimalarial combinations with chemical series in preclinical evaluation and to optimize future antimalarial therapies.

      Strengths:

      The streamlined approach relies on induction of the lacZ enzyme only after drug washout. As opposed to when stably expressed, this allows the authors to estimate parasite viability without undergoing serial dilutions to estimate viable parasite titers. This innovation vastly reduced resource and time intensity, enabling greater throughput for parasite viability estimation. The established methodology and analysis pipeline enabled the testing of 49 drug combinations for parasite viability in the MULT-i<sup>2</sup> assay compared to only 9 in the conventional cPRR assay. This provided improved resolution in the ability to estimate drug combination parameters in a pharmacodynamic model. The ability to comprehensively characterize combination pharmacodynamic properties in vitro will have important implications for downstream modelling of in vivo combinations, and for optimizing future antimalarial combination therapies.

      The authors made good use of modelling and AICc for parametric estimation and model evaluation to demonstrate the advantages of the richer dataset afforded by the MULT-i<sup>2</sup> assay.

      We thank reviewer 2 for her/his appreciation of our work.

      Weaknesses:

      The authors correctly identified a range of confounding effects that lead to artefacts in their assay results when compared to the cPRR assay. For instance, the authors observed reduced signal at high parasite density during recovery due to overgrowth and likely enzyme degradation, and suggested residual signal may remain from non-proliferating sexual stage parasites surviving drug treatment that would not be detected in the cPRR assay.

      Measurement of parasite viability in the MULT-i<sup>2</sup> assay was achieved by extrapolating the chemoluminescence signal to that of a serial dilution of parasites made at the initiation of drug treatment. How did the authors account for differing levels of enzyme expression at early (e.g., ring) vs late stage parasites (trophozoite or schizonts)? Were cultures synchronized prior to initiation of assays? Could differences in life-cycle progression following drug treatment be an additional confounding factor that may account for differences with the PRR v2 assay?

      We thank the reviewer for raising this important point. All, the MULT-i<sup>2</sup> and PRR v2 assay were performed using asynchronous parasite cultures. We will clarify this in the revised manuscript.

      We agree that parasite developmental stages may influence the MULT-i<sup>2</sup> readout, as LacZ expression levels differ between parasite stages, with differences observed between ring stages and more mature trophozoite/schizont stages as published by Hellingman et al., 2024. This represents a potential source of variability, as the MULT-i<sup>2</sup> assay quantifies the amount of expressed reporter enzyme rather than directly measuring parasite numbers at the time of readout. The use of asynchronous cultures minimizes the impact of stage-specific effects by providing a mixed parasite population representative of the natural distribution of developmental stages. Nevertheless, we acknowledge that differences in parasite stage progression following drug exposure may contribute to variation in the extrapolated parasite numbers and may partially explain differences observed between the MULT-i<sup>2</sup> and PRR v2 assay measurements. We will add this consideration to the Discussion.

      The addition of an inducible element is an improvement of their earlier lacZ/β-gal<sup>SENSOR</sup> (PMID: 41575867); however, the authors fail to explain why this is an improvement and how this adds additional merit over the initial system. While the authors compare their new assay to the PRR v2, they fail to compare it to their own non-inducible lacZ/β-gal<sup>SENSOR</sup> system. Their non-inducible system already showed superiority to the cPRR assays, and it would be good to show how they compare and what the advantages of the new system are over the old. e.g., how is the signal-to-noise improved?

      We thank the reviewer for this important comment. The main improvement provided by the inducible system is the temporal separation of parasite growth/drug exposure from reporter expression. In the original non-inducible lacZ/β-gal<sup>SENSOR</sup> system, reporter expression occurs continuously throughout the assay, resulting in accumulation of β-galactosidase during parasite growth/drug exposure and therefore an increasing background signal. Consequently, quantification relies on endpoint reporter levels and does not allow the reporter expression window to be standardized independently of parasite exposure history.

      In contrast, in the MULT-i<sup>2</sup> system, reporter expression is initiated only after addition of rapamycin post antimalarial drug washout. This prevents reporter accumulation during the drug exposure window and ensures a defined reporter enzyme accumulation window after drug exposure. Importantly, this allows parasite numbers to be extrapolated from a calibration curve generated at the time of induction, which would not be possible with the non-inducible system because reporter expression would continue after drug removal and would depend on the previous culture history.

      We will revise the manuscript to more clearly describe these advantages and to emphasize that the key benefit of the inducible system is not simply an increase in signal intensity, but improved control of reporter expression, reduced background accumulation, and the ability to perform quantitative parasite reduction rate measurements.

      How does the sensitivity compare? How quickly does the can the signal be detected after induction? They show signal after 48h, but it would be very useful to the community to look at earlier timepoints as well and compare them to the uninduced line and a line that has been induced 48h earlier to match the expression patterns throughout the lifecycle (something like 2h,4h,6h, 12h, and 24h).

      We thank the reviewer for this important suggestion. We acknowledge that the sensitivity of the MULT-i<sup>2</sup> readout depends on both the initial parasite density and the duration of the induction period and that a detailed characterization of the induction kinetics, including earlier time points after rapamycin addition, would provide additional information on the sensitivity and temporal resolution of the MULT-i<sup>2</sup> system.

      In the present study, we focused on the time window relevant for application of the assay in a PRR assay workflow and routine drug screening setting. Earlier time points (<24 h after induction) were therefore not systematically evaluated. The selected time points were chosen based on the expected kinetics of the loxP-DiCre recombination system, which has previously been reported to achieve high recombination efficiency within one asexual parasite cycle, (Collins et al., 2013) and shown with own data in this study, as well as on practical considerations for implementation in routine workflows.

      Is the chemiluminescence signal for the i-lacZ induced parasites comparable to the stably expressed lacZ parasites previously characterized by the group? If so, do the authors consider this inducible iteration a complete replacement for PRR assays?

      We thank the reviewer for this question. The chemiluminescence signal obtained with the inducible lacZ (i-lacZ) parasites is comparable to that observed with the previously characterized constitutively expressing lacZ parasites. However, the inducible system provides an important additional advantage by avoiding continuous β-galactosidase production and accumulation during parasite growth, thereby reducing background signal and enabling a controlled reporter expression window.

      We do not consider the MULT-i<sup>2</sup> assay to be a replacement for classical PRR assays. Rather, we consider it a complementary approach that enables more efficient screening and characterization of drug combinations, particularly by providing information on the time-dependent onset of parasiticidal activity in a higher-throughput format. Promising combinations identified using MULT-i<sup>2</sup> assay can subsequently be investigated in more extensive PRR assays.

      Reviewer 3 (Public review):

      In this manuscript, the authors strived to develop a highly efficient drug survival assay for in vitro cultured human malaria parasites P. falciparum. This was done by generating a transgenic P. falciparum line using a creLox strategy that allows detection of (presumably) viable parasites by a β-lactamase assay. To estimate the Limit of quantification of the recombined P. falciparum NF54i-lacZ, the authors ultimately designed a protocol in which viable parasites are detected by the luminescence of β-D-galactoside generated by β-lactamase within the transgenic parasites. For this, the parasite must be incubated with rapamycin for 120 hours to induce CreLox recombinase, which places β-lactamase under an active promoter. Using this assay, termed MULT-i<sup>2</sup>, the author shows interactions between two antimalarial drug pairs that were previously demonstrated by another assay. In the case of pyronaridine and piperaquine pair, the NULT-i<sup>2</sup> assay generated some additional insights compared to the previous assay, presumably by virtue of including more concentration datapoints. In conclusion, the authors argue that the MULT-i<sup>2</sup> assay is much less resource-intensive and time-consuming and can be applied on a large scale at a much lower cost and with the highest efficiency.

      Overall, the data generated in this manuscript are clear and well represented, and I am convinced that MULT-i<sup>2</sup> provides yet another of many drug assays for malaria parasites and could be put to good use. However, I struggle to fully appreciate the merit of his study, as the manuscript reads more like a technical document than a scientific study.

      We thank reviewer 3 for her/his appreciation of our work.

      I particularly lack an understanding of the strengths and weaknesses/limitations of the MULT-i<sup>2</sup> methodology and, thus, its applicability. I also do not fully appreciate the need for such an elaborate luminescence-based experimental setup. It would be good if some of these issues were addressed.

      Specifically:

      (1): The whole assay is based on detecting parasites by luminescence after 120 hr (5 days) after drug exposure. During that time, presumably the parasites that survived the drug pressure regrow to a detectable level and, at the same time, perform efficacious CreLox-based recombination to produce β-D-galactoside for detection. Is this necessary? How superior is this detection method to other methods, such as Fluorescence-assisted Cell Sorting (FACS), etc? Moreover, the 5-day growth-CreLox-β-D-galactoside production could introduce a series of confounding effects. In my view, more studies (beyond comparisons with a single existing method) would be useful for understanding this entire process.

      We thank the reviewer for raising this important point regarding the rationale, applicability, and limitations of the MULT-i<sup>2</sup> methodology.

      Quantification of viable parasites after drug exposure remains challenging, particularly when surviving parasites are present at low frequencies or require extended recovery periods. Current approaches, such as the parasite reduction ratio (PRR) assay based on [<sup>3</sup>H]-hypoxanthine incorporation, provide sensitive measurements of replicating parasites but are labor-intensive, require specialized infrastructure, and are not easily scalable for large numbers of drug combinations. Alternative approaches based on HRP2 detection no longer rely on radioactive readouts but generally provide lower sensitivity, particularly when quantifying low levels of surviving parasites within a shorter time frame.

      The MULT-i<sup>2</sup> assay was developed to address these limitations by combining a highly sensitive chemiluminescent β-galactosidase readout with an inducible reporter system. The 5-day induction period after drug exposure serves as a controlled gene expression step, allowing surviving parasites to recover and produce sufficient reporter signal for sensitive quantification using a standard plate reader. This approach enables higher-throughput assessment of parasiticidal activity while avoiding radioactive readouts and reducing the need for labor-intensive dilution-based approaches.

      We acknowledge that the recovery and reporter expression period introduces additional biological steps compared with direct parasite detection methods and may therefore represent a potential source of variability. The MULT-i<sup>2</sup> assay is not intended to replace all existing viability measurements but rather to provide a complementary screening tool for investigating larger numbers of drug combinations. More detailed comparisons with additional detection platforms, including fluorescence-based approaches such as flow cytometry, would be valuable; however, a comprehensive comparison of all available parasite viability readouts was beyond the scope of this study. We will add more explanations to the Discussion including the strengths and limitations.

      (2) Related to that above, how would MULT-i<sup>2</sup> perform in case of drugs that do not necessarily kill all parasites, such as artemisinin? In the case of artemisinin, it is becoming evident that at least a small fraction of the parasite revives after treatment via a temporary dormancy state. This has, in fact, also been shown for other drugs such as mefloquine, pyrimethamine, etc. Would such a situation produce a range of false readings? In general, in its current state, it is hard to see what the limitations of this method are, which makes it hard to decide whether to use it for a particular application.

      We thank the reviewer for raising this important point regarding the interpretation and applicability of the MULT-i<sup>2</sup> assay. We agree that distinguishing between growth inhibition assays and viability-based assays is essential when interpreting the response to drugs that induce temporary parasite dormancy or delayed recovery.

      The MULT-i<sup>2</sup> assay was specifically developed as a viability-based approach and therefore differs fundamentally from conventional IC50 assays, which primarily measure inhibition of parasite growth during drug exposure and may not capture parasites that survive treatment through temporary growth arrest or dormancy. Similar to the PRR assay, the MULT-i<sup>2</sup> assay measures the ability of surviving parasites to recover and proliferate after drug exposure. Therefore, parasites that temporarily enter a dormant state but subsequently resume replication are expected to contribute to the measured signal rather than representing false-positive or false-negative results.

      This is illustrated by the artemisinin experiments presented in this study, where the MULT-i<sup>2</sup> assay captures the recovery of surviving parasites following treatment as it does the PRR v2 assay.

      (3) Given the stated cost and labor efficiency of MULT-i<sup>2</sup>, it is disappointing to see only two applications for two drug pairs: atovaquone/proguanil and piperquine/pyronaridine, for both of which their interactions were already known. The manuscript would benefit greatly if the authors demonstrated more drug interactions and identified (and ultimately validated) new ones. This would certainly make MULT-i<sup>2</sup> method more attractive. In particular, it would be nice to see if one could use MULT-i<sup>2</sup> for studies of triple combinations as enthusiastically suggested.

      We thank the reviewer for this valuable suggestion. We agree that demonstrating additional applications, including triple-drug combinations, would further highlight the potential of the MULT-i<sup>2</sup> assay.

      The primary aim of this study was to validate the MULT-i<sup>2</sup> methodology against the established PRR v2 assay and to demonstrate that the new platform can reproduce known parasiticidal interaction profiles while providing a more scalable workflow. For this reason, we selected well-characterized drug combinations, including atovaquone/proguanil and piperaquine/pyronaridine, which provide suitable benchmark systems for comparison with previous PRR data.

      Although evaluation of a larger number of novel combinations and triple-drug regimens would be highly valuable, generating corresponding PRR datasets for direct comparison was beyond the scope of the current study.

      (4) Throughout the manuscript, the authors claim that MULT-i<sup>2</sup> is considerably less expensive and can be done much faster than previous methods. In my view, this is not exactly a scientific argument. The cost of an assay depends heavily on the cost of reagents and labor, which are subject to market price fluctuations. The efficiency and time consumption can very much depend on laboratory organization, etc. Unless the author could specifically demonstrate where and how these assays are cheaper and faster, I suggest not discussing this.

      We thank the reviewer for this important comment. We agree that absolute assay costs can vary depending on local reagent prices, labor costs and laboratory infrastructure.

      When comparing both methods under the same laboratory conditions, the total assay duration of the MULT-i<sup>2</sup> assay is shorter than that of the PRR assay (11 days (MULT-i<sup>2</sup>) compared with approximately 21–28 days (PRR) according to published protocols). In addition, the MULT-i<sup>2</sup> assay reduces labor-intensive processing steps and enables higher-throughput measurements using a plate reader for readout. These factors contribute to reduced workload and improved scalability, independent of fluctuations in individual reagent or personnel costs.

    1. eLife Assessment

      This important study combines anatomical tracing, tissue clearing, and functional manipulations to demonstrate lateralized brainstem control of hepatic glucose metabolism and identify a site of sympathetic nerve crossover supplying the liver. The evidence supporting the anatomical organization of hepatic sympathetic innervation is compelling, and the functional studies provide solid support for a role of asymmetric sympathetic outflow in regulating glucose homeostasis. While some uncertainty remains regarding the contribution of sensory innervation and the extent to which these findings generalize beyond mice, the work provides an invaluable advance in understanding neural regulation of liver metabolism.

    2. Reviewer #2 (Public review):

      Summary:

      The manuscript by Wang and colleagues aims to determine whether hepatic glucose metabolism is differentially regulated by the left and right sides of the LPGi and to reveal decussation of hepatic sympathetic nerves.

      The authors used tissue clearing to identify sympathetic fibers in the liver lobes, then injected PRV into the hepatic lobes. Five days post-injection, PRV-labeled neurons in the LPGi, which were identified. The results indicated contralateral dominance of premotor neurons and partial innervation of more than one lobe. Then the authors activated each side of the LPGi, resulting in a greater increase in blood glucose levels after right-sided activation than after left-sided activation, and in changes in protein expression in the liver lobes. These data suggested lobe-specific modulation of HGP. Chemical denervation of a particular lobe did not affect glucose levels due to compensation by the other lobes. In addition, nerve bundles decussate in the hepatic portal region.

      Strengths:

      The manuscript is timely and relevant. It is important to understand the sympathetic regulation of the liver and the contribution of each lobe to hepatic glucose production. The authors use state-of-the-art methodology.

      Weaknesses:

      (1) Image clarity was improved in some cases, but not in others. For example, Figure 3I, showing c-Fos expression, is not convincing due to the image quality and lack of orientation.

      (2) The methods section states that 8-weeks-old male mice were used in the experiments without specifying the experiments (e.g., brain injection with AAVs or PRV organ inoculation). The authors should include these details.

      (3) The authors should use the exact location of pre- and postganglionic neurons as they often refer to neurons in the sympathetic chain. Their findings should be compared with the existing literature on the location of preganglionic cells.

      (4) Figure legends should be revised and matched with the text.

    3. Reviewer #4 (Public review):

      Summary of General Strengths & Weaknesses:

      The studies here are highly informative for anatomical tracing and sympathetic nerve function in the liver in relation to glucose levels, but because they are conducted in a single species, it is challenging to translate them to humans or determine whether these neural circuits are evolutionarily conserved. Dual-labeling anatomical studies are elegant, and the addition of chemogenetic and optogenetic studies provides mechanistically informative. Denervation studies lack proper controls, and sensory innervation in the liver is overlooked.

      Specific Weaknesses - Major:

      (1) The species name should be included in the title.

      (2) Tyrosine hydroxylase was used to mark sympathetic fibers in the liver, but this marker also labels a portion of sensory fibers that need to be ruled out in whole-mount imaging data.

      (3) Chemogenetic and optogenetic data demonstrating hyperglycemia should be described in the context of prior work demonstrating liver nerve involvement in these processes. There is only a brief mention in the Discussion currently, but comparing methods and observations would be helpful.

      (4) Sympathetic denervation with 6-OHDA can drive compensatory increases in tissue sensory innervation, and this should be measured in the liver denervation studies to implicate potential crosstalk, especially given the increase in LPGi cFOS that may be due to afferent nerve activity. Compensatory sympathetic drive may not be the only culprit, though that is clearly assumed. The sensory or parasympathetic/vagal innervation of the liver is altogether ignored in this paper and could be better described in general.

      Comments on the revised version.

      Across all reviewer comments, the revised resubmission has adequately addressed all concerns.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Wang et al. reports the potential involvement of an asymmetric neurocircuit in the sympathetic control of liver glucose metabolism.

      Strengths:

      The concept that the contralateral brain-liver neurocircuit preferentially regulates each liver lobe may be interesting.

      Weaknesses:

      However, the experimental evidence presented did not support the study's central conclusion.

      We thank the reviewer for recognizing the conceptual novelty of our work and for constructive comments aimed at enhancing its rigor and clarity. In response, we carried out targeted experiments to address the points raised, including: (i) further characterization of LPGi projections to vagal and sympathetic circuits; (ii) evaluation of potential pancreatic involvement; and (iii) validation of the specificity of chemogenetic activation within the proposed circuit. All new experiments, figures, and text have been incorporated, and corresponding revisions are highlighted for ease of review.

      (1) Pseudorabies virus (PRV) tracing experiment:

      The liver not only possesses sympathetic innervations but also vagal sensory innervations. The experimental setup failed to distinguish whether the PRV-labeling of LPGi (Lateral Paragigantocellular Nucleus) is derived from sympathetic or vagal sensory inputs to the liver.

      Thank you for raising this important point. We fully agree that the liver receives both sympathetic and vagal sensory innervation, and we acknowledge that PRV-based tracing alone does not definitively distinguish between these two pathways. This represented a limitation of the original experimental design.

      Based on established anatomical literature as well as our experimental observations, vagal sensory neuron cell bodies reside in the nodose ganglion (NG), and their central projections terminate predominantly in the nucleus of the solitary tract (NTS) (Nature. 2023;623(7986):387-396; Curr Biol. 2020;30(20):3986-3998.e5.), which is located in the dorsomedial medulla. In contrast, the LPGi, together with other sympathetic-related nuclei, is predominantly distributed in the ventral medulla (Cell Metab. 2025;37(11):2264-2279.e10; Nat Commun. 2022;13(1):5079).

      To determine whether the LPGi contains neurons that modulate the liver via vagal sensory pathways, we performed two complementary experiments.

      First, we conducted CGRP immunohistochemistry on brainstem sections, using the NTS, a well-established visceral sensory centre, as a positive control. While abundant CGRP-positive cell bodies were detected in the NTS as expected, few to no CGRP-positive cell bodies were observed in the LPGi (Figure S1G). These results strongly support that the LPGi neurons labeled in our PRV tracing predominantly belong to sympathetic efferent circuits rather than vagal sensory pathways.

      Second, to examine whether LPGi neurons send axonal projections to sensory ganglia, we injected hSyn-Cre combined with DIO-Axon-EGFP into the LPGi and examined both the dorsal root ganglia (DRG) and nodose ganglia (NG). No Axon-EGFP-positive signals were detected in either ganglion (Figures S1H-S1J), indicating that LPGi neurons do not directly innervate sensory ganglia. In other words, PRV cannot retrogradely trace to the LPGi via the NG or DRG.

      These additions have been incorporated into the revised manuscript, with the Result 1 clearly documenting that these findings confirm that the LPGi specifically regulates sympathetic, rather than vagal sensory, inputs to the liver.

      (2) Impact on pancreas:

      The celiac ganglia not only provide sympathetic innervations to the liver but also to the pancreas, the central endocrine organ for glucose metabolism. The chemogenetic manipulation of LPGi failed to consider a direct impact on the secretion of insulin and glucagon from the pancreas.

      Thank you for this important comment. We agree that the celiac ganglia (CG) provide sympathetic innervation not only to the liver but also to the pancreas, which plays a central role in glucose homeostasis through the secretion of both insulin and glucagon. Therefore, the potential pancreatic implications associated with LPGi chemogenetic manipulation are worth careful consideration.

      To address this concern, we measured circulating glucagon and insulin levels following chemogenetic manipulation of the LPGi<sup>GAD1</sup> neurons. We found that neither glucagon nor insulin levels changed significantly under our experimental conditions, which indicated that the hyperglycemic effect induced by LPGi activation is unlikely to be mediated by changes in pancreatic hormone secretion (Figure S2G).

      These additions have been incorporated into the revised manuscript, with the Result 2 clearly documenting that these findings confirm that the hyperglycemic effect induced by LPGi activation is unlikely to be mediated indirectly via altered pancreatic endocrine output.

      (3) Neuroanatomy of the brain-liver neurocircuit:

      The current study and its conclusion are based on a speculative brain-liver sympathetic circuit without the necessary anatomical information downstream of LPGi.

      Thank you for raising this important point. A clear anatomical definition of the downstream pathways linking the brain to the liver was essential for interpreting the proposed brain-liver sympathetic circuit.

      The present study (Figure 4A) provides direct anatomical evidence supporting the organization of the brain–liver sympathetic neurocircuit. These observations are consistent with our recent detailed characterization of the brain-liver sympathetic circuit published in Cell Metabolism (Cell Metab. 2025;37(11):2264–2279). In that study, we showed that LPGi GABAergic neurons inhibit GABAergic neurons in the caudal ventrolateral medulla (CVLM). Disinhibition of CVLM reduced GABAergic suppression of rostral ventrolateral medulla (RVLM) neurons, which are key excitatory drivers of sympathetic tone. RVLM neurons project to sympathetic preganglionic neurons in the sympathetic chain (Syc). These neurons synapse with postganglionic sympathetic neurons in ganglia such as the celiac-superior mesenteric ganglion (CG-SMG). Postganglionic sympathetic fibers then innervate the liver, releasing norepinephrine (NE) to activate hepatic β<sub>2</sub>-adrenergic receptors and stimulate hepatic glucose production (HGP).

      Together, these data establish a coherent anatomical basis for the proposed brain-liver sympathetic pathway and clarify the downstream organization relevant to the functional experiments presented in figure 4A and Author response image 1..

      Author response image 1.

      Tracing scheme (Left) and whole-mount imaging (Right) of PRV-labeled brain-liver neurocircuit. Scale bars, 3,000 (whole mount) or 1,000 (optical sections) μm.

      (4) Local manipulation of the celiac ganglia:

      The left and right ganglia of mice are not separate from each other but rather anatomically connected. The claim that the local injection of AAV in the left or right ganglion without affecting the other side is against this basic anatomical feature.

      Thank you for raising this important anatomical point. We fully acknowledge that the left and right CG in mice are interconnected, and that unilateral viral injection could theoretically affect the contralateral side. The CG-SMG complex serves as a major sympathetic hub that regulates visceral organ functions. Recent transcriptomic, anatomical, and functional studies have revealed that the CG-SMG is not a homogeneous structure but is composed of molecularly and functionally distinct neuronal populations. These populations exhibit specialized projection patterns and regulate different aspects of gastrointestinal physiology, supporting a model of modular sympathetic control. (Nature. 2025 Jan;637(8047):895-902). Therefore, we were aware of this phenomenon during the initial stages of these experiments.

      To minimize unintended spread to the contralateral CG, we took two complementary approaches.

      First, we optimized the injection strategy by using an extremely small injection volume (100 nL per site), with a very slow infusion rate (50 nL/min), and fine glass micropipettes. With these refinements, contralateral viral spread was rarely observed.

      Second, and importantly, all animals included in the final analyses were subjected to post hoc anatomical verification. After completion of the experiments, CGs were collected, sectioned, and examined for viral expression. As shown in Supplementary Figure 5F, only mice in which viral expression was strictly confined to the targeted CG, with no detectable infection in the contralateral ganglion, were included in the presented data.

      Together, these measures ensure that our local manipulation of the intended CG produced the reported effects. We have revised the Methods section to more explicitly detail these technical precautions, and the legend for Figure S5F clearly states its role in validating injection specificity.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Wang and colleagues aims to determine whether the left and right LPGi differentially regulate hepatic glucose metabolism and to reveal decussation of hepatic sympathetic nerves.

      The authors used tissue clearing to identify sympathetic fibers in the liver lobes, then injected PRV into the hepatic lobes. Five days post-injection, PRV-labeled neurons in the LPGi were identified. The results indicated contralateral dominance of premotor neurons and partial innervation of more than one lobe. Then the authors activated each side of the LPGi, resulting in a greater increase in blood glucose levels after right-sided activation than after left-sided activation, as well as changes in protein expression in the liver lobes. These data suggested modulation of HGP (hepatic glucose production) in a lobe-specific manner. Chemical denervation of a particular lobe did not affect glucose levels due to compensation by the other lobes. In addition, nerve bundles decussate in the hepatic portal region.

      We thank the reviewer for the thorough and constructive evaluation of our manuscript. In direct response, we undertook comprehensive revisions to enhance the rigor and clarity of the study, including: (i) correcting ambiguous or misleading terminology about anatomical resolution and sympathetic circuit organization; (ii) expanding the Methods section with complete experimental details, improved image presentation, and explicit justification of our viral and genetic approaches; and (iii) strengthening data interpretation by addressing issues related to sparse PRV labeling, projection heterogeneity, and the functional implications of double-labeled neurons.

      Strengths:

      The manuscript is timely and relevant. It is important to understand the sympathetic regulation of the liver and the contribution of each lobe to hepatic glucose production. The authors use state-of-the-art methodology.

      Weaknesses:

      (1) The wording/terminology used in the manuscript is misleading, and it is not used in the proper context. For instance, the goal of the study is "to investigate whether cerebral hemispheres differentially regulate hepatic glucose metabolism..." (see abstract); however, the authors focus on the brainstem (a single structure without hemispheres). Similarly, symmetric is not the best word for the projections.

      We thank the reviewer for raising these critical points regarding terminology and conceptual framing. We acknowledge that certain phrases in our original manuscript may have been overly broad or ambiguous, particularly in describing the scope of sympathetic heterogeneity and the specificity of neural projections. Due to practical constraints and the scope of our study, our investigation focused on the brainstem, which represents the final common pathway for these lateralized commands. We acknowledge that terms referring to the cerebral hemispheres do not accurately describe our study. We have revised the manuscript to ensure accurate and consistent terminology.

      Below are specific examples:

      Original 1: This study aims to investigate whether cerebral hemispheres differentially regulate hepatic glucose metabolism and localize the site of sympathetic crossover to the liver.

      Revised 1: “This study aimed to determine whether the central nervous system exerts lateralized control over hepatic glucose metabolism and to localize the site of peripheral sympathetic crossover to the liver.”

      Original 2: These findings demonstrate that the brain exerts lobe-specific, lateralized control of hepatic glucose metabolism via symmetric brain-liver sympathetic pathways.

      Revised 2: “These findings demonstrate that the brainstem can exert lobe-specific, lateralized control of hepatic glucose metabolism via bilaterally projecting brain–liver sympathetic pathways.”

      Original 3: The cerebral hemispheres exhibit pronounced functional asymmetry, [1,2] a phenomenon traditionally associated with cognitive and motor processes such as language, voluntary movement, and spatial navigation.[3,4]

      Revised 3: “Pronounced functional lateralization within the central nervous system (CNS) is a well-documented phenomenon, [1,2] traditionally associated with cognitive and motor processes such as language, voluntary movement, and spatial navigation [3,4].

      (2) Sparse labeling of liver-related neurons was shown in the LPGi (Figure 1). It would be ideal to have lower magnification images to show the area. Higher quality images would be necessary, as it is difficult to identify brainstem areas. The low number of labeled neurons in the LPGi after five days of inoculation is surprising. Previous findings showed extensive labeling in the ventral brainstem at four days post-inoculation (Desmoulins et al., 2025). Unfortunately, it is not possible to compare the injection paradigm/methods because the PRV inoculation is missing from the methods section. If the PRV is different from the previously published viral tracers, time-dependent studies to determine the order of neurons and the time course of infection would be necessary.

      We sincerely thank the reviewer for these detailed and constructive comments regarding the PRV tracing experiments. We fully agree that careful presentation and interpretation of the anatomical data are essential for ensuring rigor and transparency. We address each point in detail below.

      (1) Image magnification and anatomical context of LPGi labeling

      We agree that the original images did not sufficiently convey the broader anatomical context of the LPGi. Due to fluorescence quenching in previous sections, we repeated the PRV retrograde tracing experiment and performed statistical analysis. In the revised manuscript, we replaced the original panels in Figure 1 and Figure S1 with new images that include lower-magnification overviews of the brainstem, alongside higher-magnification views of the LPGi (Figure 1). These images clearly delineate the LPGi with respect to established anatomical landmarks and atlas boundaries. Image contrast and resolution were optimized to allow unambiguous identification of PRV-labeled neurons and surrounding structures.

      (2) Sparse LPGi labeling at 5 days post-injection and methodological details

      We apologize for the omission of the detailed PRV injection protocol in the original Methods section. We deliberately used small-volume, local injections (1 µL per liver lobe) to minimize viral spread and to restrict labeling to circuits specifically connected to the targeted hepatic region. This sparse labeling was consistent with the use of small, spatially restricted injections designed to minimize off-target spread and preferentially label higher-order upstream neurons. This information has now been added, including the PRV strain, viral titer, injection volume, precise injection coordinates, and surgical procedures. All new figures, legends, and Method details have been incorporated, with changes clearly highlighted for ease of review.

      These additions have been incorporated into the revised manuscript, in Figure 1B-D, Figure S1C-F. Furthermore, we also added details of the Methods.

      (3) Not all LPGi cells are liver-related. Was the entire LPGi population stimulated, or was it done in a cell-type-specific manner? What was the strain, sex, and age of the mice? What was the rationale for using the particular viral constructs?

      We thank the reviewer for this insightful and important question. We agree that not all neurons within the LPGi are liver-related, and we apologize that our rationale was not clearly articulated in the original manuscript.

      (1) Our decision to target GABAergic neurons in the LPGi using GAD1-Cre mice was based on prior experimental evidence rather than an assumption about the entire LPGi population. In our previous study (Cell Metab. 2025;37(11):2264-2279.e10), we performed single-cell RNA sequencing on retrogradely labeled LPGi neurons following liver tracing. These analyses revealed that the majority of liver-projecting LPGi neurons are GABAergic in nature. Based on these findings, we chose to selectively manipulate GABAergic neurons in the LPGi rather than the entire LPGi neuronal population, to achieve greater cellular specificity and to minimize potential confounding effects arising from heterogeneous neuron types within this region. We regret that this rationale was not clearly described in the original submission and have now revised the manuscript to explicitly state this reasoning (Results section 2, paragraph 2: “Prior single-nucleus RNA sequencing and immunofluorescence analyses demonstrated that liver-projecting LPGi neurons are predominantly GABAergic.”).

      (2) In addition, we apologize for the omission of mouse strain, sex, and age information in the Methods section. These details have been fully added.

      (3) We selected AAV-based viral vectors, specifically the AAV9 serotype, due to their well-established efficiency in transducing neurons in the brainstem, relatively low toxicity, and widespread use in circuit-level chemogenetic and optogenetic studies. When combined with Cre-dependent viral constructs in GAD1-Cre mice, this approach enabled selective and reliable manipulation of LPGi GABAergic neurons.

      (4) The authors should consider the effect of stimulation of double-labeled neurons (innervating more than one lobe) and potential confounding effects regarding other physiological functions.

      We thank the reviewer for raising this important point. We agree that neurons innervating more than one liver lobe could, in principle, introduce potential confounding effects and may reflect higher-order integrative autonomic neurons.

      This consideration is consistent with a key finding of the cited study: the CG-SMG contains molecularly distinct sympathetic neuron populations (e.g., RXFP1<sup>+</sup> vs. SHOX2<sup>+</sup>) that exhibit complementary organ projections and separate, non‑overlapping functions. Specifically, RXFP1<sup>+</sup> neurons innervate secretory organs (pancreas, bile duct) to regulate secretion, while SHOX2<sup>+</sup> neurons innervate the gastrointestinal tract to control motility. This functional segregation supports the concept of specialized autonomic modules rather than a uniform, “fight-or-flight” response, reinforcing the need for careful interpretation of circuit-specific manipulations. (Nature. 2025;637(8047):895-902; Neuron. 2026;114(3):463-478.e7).

      In our PRV tracing experiments, the proportion of double-labeled neurons was relatively small, suggesting that the majority of labeled LPGi neurons preferentially associate with individual hepatic lobes. Nevertheless, we recognize that activation of this minority population could contribute to broader physiological effects beyond strictly lobe-specific regulation. We have therefore added a paragraph in the second paragraph of the Discussion (Paragraph 2: “A small subset of LPGi neurons was double-labeled after bilateral PRV injections, suggesting a fraction of these neurons projects bilaterally to both sides of the liver. Such neurons may support interlobar coordination.”).

      (5) The authors state that "central projections directly descend along the sympathetic chain to the celiac-superior mesenteric ganglia". What they mean is unclear. Do the authors refer to pre-ganglionic neurons or premotor neurons? How does it fit with the previous literature?

      We thank the reviewer for pointing out this imprecise wording. We agree that the original phrasing was anatomically inaccurate and potentially confusing. The pathways we intended to describe involve brainstem premotor neurons that project to sympathetic preganglionic neurons in the spinal cord. These preganglionic neurons then innervate neurons in the CG-SMG, which in turn provide postganglionic input to the liver.

      We have revised the manuscript to clearly distinguish premotor from preganglionic neurons (Results section 4, paragraph 1: “Using whole-mount clearing, we visualized the brain–liver sympathetic circuit and found that preganglionic neurons in the spinal cord send descending fibers through the sympathetic chain (SyC) to innervate postganglionic neurons in the CG-SMG (Figure 4A). Further whole-mount TH immunostaining showed that TH-positive sympathetic cell bodies within the CG-SMG project to the liver along the hepatic vasculature (Figure 4B and Figure S5F). These observations suggest that the nerve bundles likely decussate at the porta hepatis before entering the individual hepatic lobes.”).

      (6) How was the chemical denervation completed for the individual lobes?

      We thank the reviewer for raising this important methodological concern. We agree that potential diffusion of 6-OHDA is a critical issue when performing lobe-specific chemical denervation, and we apologize that our original description did not sufficiently clarify how this was controlled.

      In the revised Methods section, we provided a detailed description of the denervation procedure, including the injection volume and concentration of 6-OHDA, as well as the physical separation and isolation of individual hepatic lobes during application to minimize diffusion to adjacent tissue.

      To directly assess the specificity of the chemical denervation, we included immunofluorescence and Western blot analyses demonstrating a selective reduction of sympathetic markers in the targeted lobe (Figure 3C), with minimal effects on non-targeted lobes. These results support the effectiveness and relative spatial confinement of the 6-OHDA treatment under our experimental conditions.

      We thank the reviewer for highlighting this point, which has helped us improve both the clarity and rigor of the manuscript.

      (7) The Western Blot images look like they are from different blots, but there are no details provided regarding protein amount (loading) or housekeeping. What was the reason to switch beta-actin and alpha-tubulin? In Figures 3F -G, the GS expression is not a good representative image. Were chemiluminescence or fluorescence antibodies used? Were the membranes reused?

      We thank the reviewer for this careful and detailed evaluation of the Western blot data. We apologize that insufficient methodological detail was provided in the original submission.

      (1) We would like to clarify that the protein bands shown within each panel were derived from the same membrane. To improve transparency, we provided full, uncropped images of the corresponding membranes in the supplementary materials. In addition, detailed information regarding protein loading amounts, gel conditions, and housekeeping controls has also been added to the Methods section.

      (2) The use of different loading controls (β-actin or α-tubulin) reflects a technical consideration rather than an experimental inconsistency. In our experiments, the molecular weight of the TH (62kDa) was too close to that of α-tubulin (55kDa), and β-actin (42kDa) was therefore used to avoid band overlap and to ensure accurate quantification.

      (3) Regarding the GS signal shown in Figures 3F–G, we agree that the original representative image was suboptimal. This appears to be related to antibody performance rather than sample quality. To address this, we repeated the Western blot from Figures 3F–G using a newly validated antibody. The original tissue samples had been aliquoted and stored at −80 °C, allowing reliable re-analysis.

      (4) All Western blot experiments were detected using chemiluminescence, and membrane stripping and reprobing procedures are now explicitly described in the Methods section.

      We thank the reviewer for highlighting these issues, which significantly improve the rigor and clarity of our data presentation. All new figures and legends have been incorporated, with changes clearly highlighted for ease of review.

      (8) Key references using PRV for liver innervation studies are missing (Stanley et al, 2010 [PMID: 20351287]; Torres et al., 2021 [PMID: 34231420]; Desmoulins et al., 2025 [PMID: 39647176]).

      We thank the reviewer for pointing out these important and highly relevant references that were inadvertently omitted in our initial submission. The studies by Stanley et al. (Proc Natl Acad Sci U S A, 2010), Torres et al. (Am J Physiol Regul Integr Comp Physiol, 2021), and Desmoulins et al. (Auton Neurosci, 2025) represent key PRV-based retrograde tracing work that has mapped central neural circuits innervating the liver and thus provide essential context for our anatomical analyses.

      We agree that the inclusion of these studies is necessary to properly situate our findings within the existing literature. Accordingly, we incorporated citations to these references in the revised manuscript and discussed their relationship to our results.

      Reviewer #3 (Public review):

      Summary:

      This study found a lobe-specific, lateralized control of hepatic glucose metabolism by the brain and provides anatomical evidence for sympathetic crossover at the porta hepatis. The findings are particularly insightful to the researchers in the field of liver metabolism, regeneration, and tumors.

      Strengths:

      Increasing evidence suggests spatial heterogeneity of the liver across many aspects of metabolism and regenerative capacity. The current study has provided interesting findings: neuronal innervation of the liver also shows anatomical differences across lobes. The findings could be particularly useful for understanding liver pathophysiology and treatment, such as metabolic interventions or transplantation.

      Weaknesses:

      Inclusion of detailed method and Discussion:

      We sincerely thank the reviewer for the positive and constructive feedback, which significantly enhances both the methodological rigor and the broader biological interpretation of our study. In direct response, we revised the Discussion to elaborate on the potential physiological advantages of a lateralized and lobe-specific pattern of liver innervation. Furthermore, we expanded the Methods section to include a comprehensive description of the quantitative analysis applied to PRV-labeled neurons. Together, these revisions strengthened the manuscript’s clarity, depth, and relevance to researchers in hepatic metabolism, regeneration, and disease.

      (1) The quantitative results of PRV-labeled neurons are presented, and please include the specific quantitative methods.

      We thank the reviewer for this helpful suggestion. We have added a detailed description of the quantitative methods used to analyze PRV-labeled neurons in the revised Methods section. We have now provided detailed information in the Methods section, including the criteria used for cell counting, the anatomical boundaries of the brain regions analyzed, the delineation of regions of interest, and the normalization procedures applied to derive the reported neuron counts. These additions have been incorporated into the revised Methods, with all changes clearly indicated for ease of review.

      (2) The Discussion can be expanded to include potential biological advantages of this complex lateralized innervation pattern.

      We appreciated the reviewer’s suggestion. We have expanded the Discussion to include a paragraph addressing the potential biological significance of lateralized liver innervation. We highlight that this asymmetric organization could allow for more precise, lobe-specific regulation of hepatic metabolism, enable integration of distinct physiological signals, and potentially provide robustness against perturbations. The additional discussion content has been highlighted in the revised version as indicated (Discussion section, paragraph 3: “Bilateral LPGi activation produced additive effects, indicating that both sides of the brainstem can cooperatively regulate hepatic metabolism in a spatially segregated manner. This pattern suggests that hepatic glucose output can be modulated in a lobe-specific, rather than uniform whole-organ, manner.”).

      Reviewer #4 (Public review):

      Summary:

      The studies here are highly informative in terms of anatomical tracing and sympathetic nerve function in the liver related to glucose levels, but given that they are performed in a single species, it is challenging to translated them to humans, or to determine whether these neural circuits are evolutionarily conserved. Dual-labeling anatomical studies are elegant, and the addition of chemogenetic and optogenetic studies is mechanistically informative. Denervation studies lack appropriate controls, and the role of sensory innervation in the liver is overlooked.

      We sincerely appreciate the reviewer's thoughtful evaluation and fully agree that findings derived from a single-species model must be interpreted with caution in relation to human physiology. In direct response, we revised the manuscript to explicitly clarify that all experimental data were obtained in mice and to provide a discussion of the limitations regarding direct extrapolation to humans. Concurrently, we expanded the Discussion section by integrating our findings with recent human and translational studies, including a multicenter clinical trial demonstrating that catheter-based endovascular denervation of the celiac and hepatic arteries significantly improved glycemic control in patients with poorly controlled type 2 diabetes, without major adverse events (Signal Transduct Target Ther. 2025;10(1):371). While our current work focuses on defining the anatomical organization and functional asymmetry of this circuit in mice, the clinical findings suggest that the core principles, sympathetic control of hepatic glucose metabolism via CG-liver pathways, may be conserved and of translational relevance. Additionally, we clarified the interpretation of TH labeling and expanded the discussion of hepatic sensory and parasympathetic innervation, acknowledging their important roles in liver-brain communication and identifying them as key directions for future research. Collectively, these revisions provide a more balanced, clinically informed, and rigorous framework for interpreting our findings.

      Specific Weaknesses - Major:

      (1) The species name should be included in the title.

      We thank the reviewer for this suggestion. We agree that the species should be clearly indicated. The findings presented in this study were obtained in mice using tissue clearing and whole-organ imaging approaches. Due to technical limitations, these observations are currently restricted to the mouse strain. We have updated the title (Symmetric brain-liver circuits mediate lateralized regulation of hepatic glucose output in mice) and clarified the species used throughout the manuscript.

      (2) Tyrosine hydroxylase was used to mark sympathetic fibers in the liver, but this marker also hits a portion of sensory fibers that need to be ruled out in whole-mount imaging data

      We thank the reviewer for pointing this out. We acknowledge that TH labels not only sympathetic fibers but also a subset of sensory fibers. We have added a limitation of this point in the revised manuscript. In addition, using SyGlass (2.4.0) three-dimensional reconstruction, we observed TH-positive nerve fibers originating from the CG-SMG extending along the porta hepatis and penetrating into the liver parenchyma. Given that the CG-SMG is a well-established sympathetic ganglion innervating visceral organs (Nature. 2025 Jan;637(8047):895-902.), these nerve fibers can be definitively identified as sympathetic. In parallel, we collected DRG from spinal segments T1-6 and T7-12 five days after intrahepatic PRV injection. While T7-12 DRG are known to contain sensory neurons innervating the liver, only a sparse number of PRV-positive neurons were detected in these segments (Anat Rec A Discov Mol Cell Evol Biol. 2004 Sep;280(1):827-35. Auton Neurosci. 2024 Jun;253:103174). The additional figure and discussion content have been highlighted in the revised version as indicated (Discussion section, paragraph 6: “Third, although whole-mount TH immunostaining with three-dimensional reconstruction revealed sympathetic nerve bundles projecting from the CG to the liver, TH is not entirely specific and can also label a subset of sensory neurons. More selective approaches, such as genetic targeting of sympathetic lineages, will be important for further validation.”).

      Author response image 2.

      Representative immunofluorescence images of PRV-labeled neurons (EGFP) in DRG from the spinal segments T1-6 (bottom) and T7-12 (top) following PRV injections into the liver lobes. Scale bars, 200μm

      (3) Chemogenetic and optogenetic data demonstrating hyperglycemia should be described in the context of prior work demonstrating liver nerve involvement in these processes. There is only a brief mention in the Discussion currently, but comparing methods and observations would be helpful.

      We thank the reviewer for this suggestion. Previous studies largely relied on electrical stimulation to modulate liver innervation, which provides relatively coarse control of neural activity (Eur J Biochem. 1992;207(2):399-411). By contrast, our use of chemogenetic and optogenetic approaches allows selective, cell-type-specific manipulation of LPGi neurons. We revised the Discussion to place our functional data in the context of prior work, highlighting how these more precise approaches improve understanding of the contribution of liver-innervating neurons to hyperglycemia. The newly added discussion has been clearly labeled in the response to facilitate your review (Discussion section, paragraph 3: “This spatial organization is likely obscured by conventional electrical stimulation, which indiscriminately activates heterogeneous sympathetic fibers. By contrast, chemogenetic and optogenetic approaches permit selective, cell type-specific manipulation of LPGi neurons, thereby revealing the contralateral and lobe-specific architecture of brain-liver sympathetic control”).

      (4) Sympathetic denervation with 6-OHDA can drive compensatory increases to tissue sensory innervation, and this should be measured in the liver denervation studies to implicate potential crosstalk, especially given the increase in LPGi cFOS that may be due to afferent nerve activity. Compensatory sympathetic drive may not be the only culprit, though it is clearly assumed to be. The sensory or parasympathetic/vagal innervation of the liver is altogether ignored in this paper and could be better described in general.

      We thank the reviewer for this insightful and important comment, which highlights a potential alternative interpretation of our findings. We agree that chemical sympathetic denervation with 6-OHDA may induce compensatory changes in non-sympathetic inputs, including sensory and parasympathetic (vagal) innervation of the liver.

      Conceptually, we agree with the reviewer’s perspective that the central nervous system operates as a highly integrated homeostatic regulatory system, continuously receiving and integrating a broad range of afferent signals. These inputs include, as noted by the reviewer, hepatic sensory and vagal afferents (Science. 2024;386(6722):673-677), as well as centrally derived interoceptive signals such as brain glucose, temperature sensing, even the pulsation of cerebral vascular system (Cell Metab. 2025;37(11):2264-2279.e10.; Cell Metab. 2022;34(6):888-901.e5; Science. 2024;383(6682):eadk8511). The CNS integrates these diverse signals and generates coordinated efferent outputs to maintain systemic homeostasis.

      From this viewpoint, the changes in c-FOS activity that we observe in the LPGi likely represent only a limited snapshot of this broader integrative process, rather than evidence of a single dominant pathway. We acknowledge that compensatory sensory or parasympathetic mechanisms, in addition to altered sympathetic drive, contributed to the observed LPGi activation following hepatic sympathetic denervation.

      We further acknowledge that, due to limitations in scope and experimental focus, we did not directly assess sensory or parasympathetic innervation of the liver in the present study. As appropriately pointed out by the reviewer, a more comprehensive characterization of hepatic neural inputs would provide a more complete picture of the underlying neurocircuitry. To address this, we expanded the Discussion and explicitly noted this limitation, including a more balanced discussion of potential crosstalk among sympathetic, sensory, and parasympathetic pathways and how these may collectively influence LPGi activity. For your convenience, the newly added discussion text has been distinctly marked in the manuscript (Discussion section, paragraph 4: “Although enhanced sympathetic output appears to mediate much of this compensation, our findings suggest that the underlying regulation extends beyond a purely descending pathway. In particular, c-FOS activation in the contralateral LPGi after unilateral 6-OHDA-mediated denervation suggests that the loss of peripheral input may be sensed through an ascending neural pathway, centrally integrated, and translated into compensatory sympathetic output to the intact hepatic lobes. These results therefore support a model in which hepatic glucose production is regulated by an integrated afferent-central-efferent loop, with our current analyses primarily resolving its efferent component.”).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Although the findings are interesting, this reviewer has major concerns about the experimental design, methodology, results, and interpretation of the data. Experimental details are lacking, including basic information (age, sex, strain of mice, procedures, magnification, etc.).

      We thank the reviewer for this important recommendation. We agree that comprehensive reporting of experimental details is essential for rigor and reproducibility.

      In the revised manuscript, we added complete information regarding mouse strain, sex, age, and sample size for each experiment. In addition, detailed descriptions of surgical procedures, viral constructs, injection parameters, imaging magnification, and analysis methods have been incorporated into the Methods section.

      These revisions ensured that all experiments are described with sufficient technical detail and clarity to allow accurate interpretation and replication of our findings. Experimental details have been incorporated, and corresponding revisions are highlighted for ease of review.

      Reviewer #3 (Recommendations for the authors):

      Addressing a few questions might help:

      (1) The study found that liver-associated LPGi neurons are predominantly GABAergic. It would be informative to molecularly characterize the PRV-traced, liver-projecting LPGi neurons to determine their neurochemical phenotypes.

      We thank the reviewer for this insightful suggestion. We agree that molecular characterization of liver-projecting LPGi neurons is important for understanding their functional identity.

      This issue has been addressed in detail in our recent study (Cell Metab. 2025;37(11):2264-2279.e10), in which we performed single-cell RNA sequencing on retrogradely traced LPGi neurons connected to the liver. These analyses demonstrated that the majority of liver-projecting LPGi neurons are GABAergic, with a defined transcriptional profile distinct from neighboring non–liver-related populations.

      Based on these findings, the current study selectively targeted GABAergic LPGi neurons using GAD1-Cre mice. We have explicitly cited these molecular results in the revised manuscript to clarify the neurochemical identity of the PRV-traced LPGi neurons. New text has been incorporated, and corresponding revisions are highlighted for ease of review.

      (2) Is it possible to do a local microinjection of a sodium channel blocker (e.g., lidocaine) or an adrenergic receptor antagonist into the porta hepatis? That would potentially provide additional evidence for the porta hepatis as the functional crossover point.

      We appreciated the reviewer’s thoughtful suggestion. Although pharmacological blockade at the porta hepatis can modulate local neural activity, this approach is inherently limited in its ability to distinguish between ipsilateral and contralateral inputs. Consequently, it may not provide definitive evidence for neural crossover at this specific site.

      In our view, the anatomical evidence provided by whole-mount tissue clearing, dual-labeled tracing, and direct visualization of decussating nerve bundles at the porta hepatis offers a more definitive demonstration of sympathetic crossover. Pharmacological blockade would affect both crossed and uncrossed fibers simultaneously and therefore would not specifically resolve the anatomical organization of this decussation.

      Nevertheless, we agree that functional interrogation of the porta hepatis represents an interesting direction for future work, and we acknowledge this possibility in the Discussion (Paragraph 6: “Fourth, although our data support a peripheral decussation at the porta hepatis, direct validation of this crossover site was not feasible with local pharmacological blockade, as currently available approaches lack sufficient spatial specificity and would likely perturb multiple neural components. Future studies employing more selective inhibitory strategies will be required to directly test this possibility.”).

      (3) It is possible to investigate the effects of unilateral LPGi manipulation or ablation of one side of CG/SMG on liver metabolism, such as hyperglycemia?

      We thank the reviewer for this important suggestion. Because unilateral LPGi manipulation was already examined in our study (Figure 2D), we focused here on unilateral ablation of the CG to further assess lateralized sympathetic control of hepatic metabolism. We successfully performed unilateral CG ablation without LPGi manipulation, but observed no significant change in blood glucose compared with the sham group (Author response image 3A and 3B). To determine whether glucose homeostasis was nonetheless affected, we further performed glucose tolerance tests (GTT) and insulin tolerance tests (ITT) (Author response image 3C and 3D). Neither test showed significant impairment after unilateral ablation, suggesting that compensatory neural mechanisms and/or hormonal homeostatic regulation may be recruited to preserve systemic glucose homeostasis.

      Author response image 3.

      (A) Blood glucose levels in mice subjected to left- or right-sided CG ablation via 6-OHDA treatment (n = 6). (B) Representative images of ablation of CG. Scale bars, 100 μm. (C and D) Blood glucose levels during GTT (C, n = 6) and ITT (D, n = 6) in mice with left- or right-sided CG ablation.

      Reviewer #4 (Recommendations for the authors):

      In the abstract and elsewhere, the use of the term 'sympathetic release' is unclear - do you mean release of nerve products, such as the neurotransmitter norepinephrine? This should be more clearly defined.

      We thank the reviewer for pointing out this ambiguity. We agree that the term “sympathetic release” was imprecise. In the revised manuscript, we explicitly referred to the release of sympathetic neurotransmitters, primarily norepinephrine, from postganglionic sympathetic fibers.

      We revised the wording throughout the manuscript to ensure accurate and consistent terminology and to avoid potential confusion regarding the underlying neurobiological mechanisms.

      Original: “Following unilateral hepatic denervation, contralateral LPGi activation induced metabolic compensation in the remaining innervated lobes, characterized by increased sympathetic release, glucose production, and glycogen depletion.”

      Revised: “Following unilateral hepatic denervation, contralateral LPGi activation induced metabolic compensation in the remaining innervated lobes, characterized by increased norepinephrine release, glucose production, and glycogen depletion.”

    1. eLife Assessment

      This valuable study demonstrates that self-motion strongly affects neural responses to visual stimuli, comparing humans moving through a virtual environment to passive viewing. The evidence for visuomotor mismatch responses is solid, although the interpretation in terms of prediction remains somewhat preliminary. This study bridges human and rodent studies on the role of prediction in sensory processing, and is therefore expected to be of interest to a large community of neuroscientists.

    2. Reviewer #1 (Public review):

      In this paper, Solyga, Zelechowski & Keller study human visuomotor mismatch responses as an alternative instantiation of prediction errors to classic oddball paradigms. Using VR, they created a condition in which participants were moving around thereby creating a visuomotor coupling between physical movement and visual flow. To attempt to isolate the contribution of specifically movement-related predictions in this condition, they contrasted it to a condition in which participants were seated and rewatching their movement trajectory during the 'active' condition. Visuomotor mismatches were created by temporarily decoupling movement and visual experience by halting the VR display as participants continued to move.

      The core finding of the paper is that participants exhibit a positively-valenced response to the visuomotor decoupling in the active but not in the passive condition. Since walking speed only insignificantly slows down following decoupling events in the active conditions, the authors argue that this difference cannot be accounted for by "changes in participants' behavior or to simple visual offset responses" with the latter being equal across both conditions. The following reinstatement of the coupling in turn does not differ between the two conditions. The authors additionally show that this mismatch response differs from visual onset responses elicited by checkerboard inversions and that it's "qualitatively" stronger than more commonly studied auditory oddball mismatch responses.

      The design with its focus on ecological validity is impressive, well-rationalized and the results are well illustrated. I additionally appreciate the control analyses with regards to changes in walking speed and playback DOF and, now added, additional participants who experience the passive condition before the active.

      My main question in round 1 regarded the isolation of visuomotor mismatch. Although the comparison with a seated control seems like a very sensible way to control for simple visual responses, there seem to be more differences than just a break in visuomotor coupling between the conditions. I therefore wonder whether the reduced offset response in the seated condition may be, in part, explained differently. For example, given that participants always conduct the active condition before rewatching their movement in the seated condition, it seemed likely that there is a component of learning across the session that flow will sometimes be halted. This is confirmed with the analyses. The explanation that there is a visuomotor component here is given further weight by their conduction of an additional group of participants who perform the conditions in the reverse order, so this has strengthened the manuscript considerably. However, it does of course remain an imperfect control because the visual stimulus is now different between the conditions for these participants. It's the best that can be achieved with this type of paradigm though and of course it yields a great deal of ecological validity.

      I was also wondering whether the authors may consider the findings in frontal electrodes more closely given that the title of the paper focuses on a specifically occipital effect. Their further analyses have confirmed that there are likely interesting frontal effects. From a theoretical point of view, the spatial dissociation in adaptation effects, which were stronger in frontal and weaker in occipital areas, seems interesting and perhaps worth discussing, especially given the interpretation that "mismatch processing may initially arise in sensory visual areas before engaging higher-order frontal regions." How come the frontal decrease in responses is not accompanied by an analogous decrease in its supposed occipital source? Could these two responses reflect different kinds of prediction error signals (i.e. objective vs subjective)?

      I remain concerned that the authors fight too defensively that they have absolutely isolated visuomotor prediction mechanisms with this paradigm. It's a nice, informative study, but it seems odd to argue there are no other possible explanations. One picks a design to optimize some features, but they will always come at some cost to others. Prioritising ecological validity, which is a justifiable aim, necessarily usually weakens some control over confounds.

      To outline my reasoning fully: My concerns wrt generic influences of action on perception are reflected in Fig 1. The P1 is smaller when walking than sitting. It seems likely that the mismatch response reflects something about extrapolation or prediction, because it is larger when walking. However, it's not necessarily sensorimotor prediction. Even if you remove action from the equation, the flow can be extrapolated or predicted most of the time in a way it cannot so well when the video is halted. Of course, the sitting condition somewhat controls for it, but when it came second the visual flow disruptions were more predictable here. A reduction in effects over time is indeed confirmed with their analyses. They now have conducted a study with the conditions in the reverse order and they find the same thing. But of course, this necessitates non-identical visual flow because the sitting condition is playing the previous participant's flow. So it is likely that across all of these comparisons, it is the visuomotor mismatch that is especially salient. It's just that each comparison is a bit messy/confounded. It would strengthen the manuscript if there were some consideration given to the other processes likely at play here.

      As a more minor point in response to our previous review, whether particular accounts represent an 'orthodox' view at present does not determine whether they raise logical issues in need of consideration. The authors may have missed that the papers in question consider mechanisms underlying the attenuation of particular pieces of information *from perception*. Not perceptual processing. We have one percept at any one moment in time and must understand how different population types synergistically generate that percept.

      Similarly, a little strange is the way in which the authors aggressively defend the position that self-generated motion is 'the strongest' type of prediction. Sure, we probably experience the effects of our actions more often than ambulances. But what about objects obeying laws of gravity or others' faces being structured and moving in systematic ways? It is hard to quantify, such that presumably many scientists would be skeptical of such a claim, and it is not needed logically to justify the importance of examining mechanisms enabling action to shape perceptual processing. I'd assume it better to fight the battles you need to (and can) fight, such that the robust claims carry more weight.

      Comments on latest version.

      Nice to see the added extra analyses. Can't see any more will be achieved via further rounds and happy with the summary to stand as is.

    3. Reviewer #2 (Public review):

      Summary:

      This study investigates whether visuomotor mismatch responses can be detected in humans. By adapting paradigms from rodent studies, the authors report EEG evidence of mismatch responses during visuomotor conditions and compare them to visual-only stimulation and mismatch responses in other modalities.

      Strengths:

      - Authors use a creative experimental design to elicit visuomotor mismatch responses in humans.

      - The study provides an initial dataset and analytical framework that could support future research on human visuomotor prediction errors.

      Weaknesses:

      - Methodological issues (e.g., volume conduction) make it difficult to confidently attribute the observed mismatch responses to activity in visual cortical regions. This could be alleviated by increasing the number of channels.

      The authors successfully demonstrate that visuomotor mismatch paradigms can, in principle, be applied in human EEG. This approach provides a translational bridge between rodent and human work on predictive processing.

      Comments on latest version.

      The authors added a brief discussion paragraph which addresses my previous comment.

    4. Reviewer #3 (Public review):

      Solyga, Zelechowski, and Keller present a concise report of an innovative study demonstrating clear visuomotor mismatch responses in ambulating humans, using a mobile EEG setup and virtual reality. Human subjects walked around a virtual corridor while EEGs were recorded. Occasionally, motion and visual flow were uncoupled, and this evoked a mismatch response that was strongest in occipitally placed electrodes and had a considerable signal to noise ratio. It was robust across participants and could not be explained by the visual stimulus alone.

      This is an important extension of their prior work in mice and represents an elegant translation of those previous findings to humans, where future work can inform theories of e.g. psychiatric diseases that are believed to involve disordered predictive processing. For the most part, the authors are appropriately circumspect in their interpretations and discussions of the implications. The paper in its current form represents an important addition to the literature.

      The authors have included analyses of the auditory mismatch using temporal electrodes, referenced to Cz (and therefore should exhibit a mismatch positivity). This added data clearly and convincingly shows that the sensorimotor mismatch is, indeed, stronger than the passive auditory MMN.

      Comments on latest version:

      The authors added useful points to the discussion and also included time frequency analyses to the paper formally, which strengthens the translational potential, in addition to the bolstering their claims slightly.

    5. Author response:

      The following is the authors’ response to the previous reviews

      We thank you for the time you took to review our work and for your feedback! The main changes to the manuscript are:

      We added a paragraph to the Discussion addressing differences in visuomotor mismatch responses recorded over frontal and occipital electrodes, and their possible interpretation.

      We added time-frequency power and phase-locking analysis as supplementary figures to the manuscript.

      We added a statement in the Discussion emphasizing the importance of performing these experiments with denser EEG channel coverage.

      Public Reviews:

      Reviewer #1 (Public review):

      In this paper, Solyga, Zelechowski & Keller study human visuomotor mismatch responses as an alternative instantiation of prediction errors to classic oddball paradigms. Using VR, they created a condition in which participants were moving around thereby creating a visuomotor coupling between physical movement and visual flow. To attempt to isolate the contribution of specifically movement-related predictions in this condition, they contrasted it to a condition in which participants were seated and rewatching their movement trajectory during the 'active' condition. Visuomotor mismatches were created by temporarily decoupling movement and visual experience by halting the VR display as participants continued to move.

      The core finding of the paper is that participants exhibit a positively-valenced response to the visuomotor decoupling in the active but not in the passive condition. Since walking speed only insignificantly slows down following decoupling events in the active conditions, the authors argue that this difference can not be accounted for by "changes in participants' behavior or to simple visual offset responses" with the latter being equal across both conditions. The following reinstatement of the coupling in turn does not differ between the two conditions. The authors additionally show that this mismatch response differs from visual onset responses elicited by checkerboard inversions and that it's "qualitatively" stronger than more commonly studied auditory oddball mismatch responses.

      The design with its focus on ecological validity is impressive, well-rationalized and the results are well illustrated. I additionally appreciate the control analyses with regards to changes in walking speed and playback DOF and, now added, additional participants who experience the passive condition before the active. I have a couple of questions/comments.

      My main question in round 1 regarded the isolation of visuomotor mismatch. Although the comparison with a seated control seems like a very sensible way to control for simple visual responses, there seem to be more differences than just a break in visuomotor coupling between the conditions. I therefore wonder whether the reduced offset response in the seated condition may be, in part, explained differently. For example, given that participants always conduct the active condition before rewatching their movement in the seated condition, it seemed likely that there is a component of learning across the session that flow will sometimes be halted. This is confirmed with the analyses. The explanation that there is a visuomotor component here is given further weight by their conduction of an additional group of participants who perform the conditions in the reverse order, so this has strengthened the manuscript considerably. However, it does of course remain an imperfect control because the visual stimulus is now different between the conditions for these participants. It's the best that can be achieved with this type of paradigm though and of course it yields a great deal of ecological validity.

      The reviewer is correct. But one should keep in mind that our result here stands in the context of a considerable amount of work on mouse cortex investigating responses to very similar visuomotor mismatches. There we can we have much additional evidence to argue that the cortical response to a visuomotor mismatch is a prediction error. We would argue, it is the best one can do in human experiments.

      I was also wondering whether the authors may consider the findings in frontal electrodes more closely given that the title of the paper focuses on a specifically occipital effect. Their further analyses have confirmed that there are likely interesting frontal effects. From a theoretical point of view, the spatial dissociation in adaptation effects, which were stronger in frontal and weaker in occipital areas, seems interesting and perhaps worth discussing, especially given the interpretation that "mismatch processing may initially arise in sensory visual areas before engaging higher-order frontal regions." How come the frontal decrease in responses is not accompanied by an analogous decrease in its supposed occipital source? Could these two responses reflect different kinds of prediction error signals (i.e. objective vs subjective)?

      We have added a paragraph to the Discussion addressing the differences between signals recorded over frontal and occipital electrodes, as suggested.

      I remain concerned that the authors fight too defensively that they have absolutely isolated visuomotor prediction mechanisms with this paradigm. It's a nice, informative study, but it seems odd to argue there are no other possible explanations. One picks a design to optimize some features but they will always come at some cost to others. Prioritising ecological validity, which is a justifiable aim, necessarily usually weakens some control over confounds.

      We are not sure what the reviewer is referring to here. We certainly do not think (or are aware of having argued) that a visuomotor prediction error is the only possibly interpretation of the response. In the last paragraph of our response to the reviewers point 3 in the last revision, we explicitly discuss that the interpretation of the responses as a prediction error is only one possible interpretation. Our argument is that it is the most likely given the evidence.

      To outline my reasoning fully: My concerns wrt generic influences of action on perception are reflected in Fig 1. The P1 is smaller when walking than sitting. It seems likely that the mismatch response reflects something about extrapolation or prediction, because it is larger when walking. However, it's not necessarily sensorimotor prediction. Even if you remove action from the equation, the flow can be extrapolated or predicted most of the time in a way it cannot so well when the video is halted. Of course the sitting condition somewhat controls for it, but when it came second the visual flow disruptions were more predictable here. A reduction in effects over time is indeed confirmed with their analyses. They now have conducted a study with the conditions in the reverse order and they find the same thing. But of course this necessitates non-identical visual flow because the sitting condition is playing the previous participant's flow. So it is likely that across all of these comparisons, it is the visuomotor mismatch that is especially salient. It's just that each comparison is a bit messy/confounded. It would strengthen the manuscript if there were some consideration given to the other processes likely at play here.

      We would be happy to add additional considerations to other processes. If the reviewer has anything specific in mind, we can add that, but it would need to be somewhat concrete with some theoretical basis. We share the reviewer’s intuition, but unless this can be formalized to the point of being experimentally testable, we do not see any value in discussing it in the manuscript.

      Regarding the reason for a difference in visual responses in walking vs sitting state is, this is not entirely clear to us. Predictive processing would provide one possible explanation. Assuming the precision weighting of predictions is higher during walking, the sudden appearance of a visual stimulus might lead to stronger stimulus history prediction errors than when just sitting. But this is rather speculative.

      As a more minor point in response to our previous review, whether particular accounts represent an 'orthodox' view at present does not determine whether they raise logical issues in need of consideration. The authors may have missed that the papers in question consider mechanisms underlying the attenuation of particular pieces of information ‘from perception’. Not perceptual processing. We have one percept at any one moment in time and must understand how different population types synergistically generate that percept.

      Please excuse, the reviewer is correct, the orthodoxy of an idea is not relevant. For dubious reasons, we chose to euphemize what we actually meant to say here. With regards to circuit implementations of predictive processing (we cannot and do not intend to speak to interpretations of predictive processing that relate to conscious perception much of V1 activity is likely not consciously perceived – we assume this is what the reviewer is referring to by “we have one percept”) – the reviewers interpretation was not unorthodox, but rather incorrect (which is what we should have said). The statement that “the brain predictively ‘cancels’ expected action outcomes from perception” is incorrect in the context of sensory processing – based on both theoretical models of predictive processing, and more importantly physiological evidence. If the point was only in regards to conscious perception, we also suspect the statement is wrong, but even if it were correct, don’t see how it pertains to our work.

      Similarly a little strange is the way in which the authors aggressively defend the position that self-generated motion is 'the strongest' type of prediction. Sure, we probably experience the effects of our actions more often than ambulances. But what about objects obeying laws of gravity or others' faces being structured and moving in systematic ways? It is hard to quantify, such that presumably many scientists would be skeptical of such a claim, and it is not needed logically to justify the importance of examining mechanisms enabling action to shape perceptual processing. I'd assume it better to fight the battles you need to (and can) fight, such that the robust claims carry more weight.

      We believe it is absolutely essential for the progress of the field that we start to emphasize the differences between something that is “predictable in principle” and “predicted by the brain”. There is likely indeed a hierarchy of predictability that looks something like this:

      (1) Sensorimotor coupling

      (2) Laws of physics

      (3) Behavior of other living things

      (4) Artificial, human-made statistical relationships

      Almost all published experiments are based on the fourth type of prediction. Indeed, why not use physics simulations instead of oddballs and MMN? We absolutely should! But the field tends to revert to artificial couplings. As a direct consequence of this, the number of papers appearing recently (from both human and mouse fields), that are built on the following premise:

      (1) Expose an animal or human to an artificial coupling between A and B (e.g. an oddball, or a global oddball, or any of a myriad other constructions).

      (2) Probe for prediction error responses to the violation of the artificial coupling.

      (3) Find no prediction error responses and conclude predictive processing is wrong.

      Is utterly baffling. The fallacy here is of course the assumption that if something is predictable in principle, the brain must predict it. Thus, we are, and will continue to be strong on this point, and we think it is essential that we – as a field – are.

      Hope these comments are helpful.

      Reviewer #2 (Public review):

      Summary:

      This study investigates whether visuomotor mismatch responses can be detected in humans. By adapting paradigms from rodent studies, the authors report EEG evidence of mismatch responses during visuomotor conditions and compare them to visual-only stimulation and mismatch responses in other modalities.

      Strengths:

      Authors use a creative experimental design to elicit visuomotor mismatch responses in humans.

      The study provides an initial dataset and analytical framework that could support future research on human visuomotor prediction errors.

      Weaknesses:

      Methodological issues (e.g., volume conduction) make it difficult to confidently attribute the observed mismatch responses to activity in visual cortical regions. This could be alleviated by increasing the number of channels.

      We have added a discussion of this.

      The authors successfully demonstrate that visuomotor mismatch paradigms can, in principle, be applied in human EEG. This approach provides a translational bridge between rodent and human work on predictive processing.

      Reviewer #3 (Public review):

      Solyga, Zelechowski, and Keller present a concise report of an innovative study demonstrating clear visuomotor mismatch responses in ambulating humans, using a mobile EEG setup and virtual reality. Human subjects walked around a virtual corridor while EEGs were recorded. Occasionally, motion and visual flow were uncoupled, and this evoked a mismatch response that was strongest in occipitally placed electrodes and had a considerable signal to noise ratio. It was robust across participants and could not be explained by the visual stimulus alone.

      This is an important extension of their prior work in mice, and represents an elegant translation of those previous findings to humans, where future work can inform theories of e.g. psychiatric diseases that are believed to involve disordered predictive processing. For the most part, the authors are appropriately circumspect in their interpretations and discussions of the implications. The paper in its current form represents an important addition to the literature.

      The authors have included analyses of the auditory mismatch using temporal electrodes, referenced to Cz (and therefore should exhibit a mismatch positivity). This added data clearly and convincingly shows that the sensorimotor mismatch is, indeed, stronger than the passive auditory MMN.

      The reference electrode placed at Cz makes it is difficult to interpret relative differences between frontal and occipital electrode responses, as the occipital electrodes are placed farther away from the Cz reference than the frontal electrodes. Similarly, signal occuring cortically near the Cz reference might only appear as though it is occipitally distributed in this montage. It is common in EEG research to remontage the data to an averaged common reference in order to better interpret the scalp distributions. As the electrode coverage was sparse for some subjects, this could be challenging, and this reviewer does not feel that it is necessary to do this analysis step, or even to drastically rewrite the body of the paper. We only request that some discussion, however brief, is included in the discussion section or the methods that recommend more dense electrode coverage in the future to better interpret scalp distributions and potential meso-scale sources.

      We have added a discussion of this as suggested.

      This is just a suggestion. The authors are encouraged to analyse (and report) time-frequency power and phase locking for these mismatch responses, as is common in much of the literature (see Roach et al 2008 Schizophrenia Bulletin). This is not to say that doing so will yield insights into oscillations per se, but converting the data to the time-frequency domain provides another perspective that has some advantages. fosters translations to rodent models, as ERP peaks do not map well between species, but e.g. delta-theta power does (see Lee et al 2018 Neuropsychopharmacology; Javitt et all 2018 Schizophrenia research; Gallimore et al 2023 Cereb Ctx). Further, ERP peaks can be influenced by the actual neuroanatomy of an individual (especially for quantifying V1 responses). Time frequency analyses may aid in interpreting the "early negative deflection with a peak latency of 48 ms " finding as well. As it stands, the report is complete, and it would be acceptable if the authors chose to save this type of analysis for a future publication.

      We have added this as suggested.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      The authors have addressed most of my concerns by providing additional analyses, partly based on new data. The volume conduction issue is partly addressed based on the result showing latency differences, however to confidently assign responses to visual regions, one would need to perform recordings with a larger number of electrodes, sufficient to perform source localization. Nevertheless, the manuscript is now more solid than the previous version.

      We have now added this point to the Discussion.

      Reviewer #3 (Recommendations for the authors):

      The reviewer appreciates that the authors have carried out time-frequency analyses, and are ok with them leaving this out of this paper.

      We have now added this to the manuscript.

      Finally, in response to the participant quote "are you printing this? hi mom!" - this reviewer concedes that it does not significantly detract from the report, and, in the interest of amusement and joy, would abide its reinstatement.

      We greatly appreciate the reviewers entertaining our attempts at humor but will leave it out as originally suggested.

    1. eLife Assessment

      This study presents a key finding: self-generated mechanical stresses enable collective protocell proliferation without dedicated division machinery, offering insight into primitive life's population growth. While quantitative imaging, membrane tension measurements, and computational modeling support the mechanism, establishing causal links between deformation and division and testing sensitivity assumptions would strengthen the work. Overall, the work reports important findings, and although the evidence in support of the conclusions is largely solid, some incomplete elements need to be addressed.

    2. Reviewer #1 (Public review):

      Li and Wu, in this article, explore the proliferation of wall-less L-forms derived from Bacillus subtilis as mimics for protocells and report an interesting new mechanism for their proliferation. The authors carry out live-cell imaging of the L-forms and find that the clusters of cells forming proto-colonies proliferate better than the isolated single cells of L-forms. They further examine the causes for this indefinite proliferation of proto-colonies of L-forms, as compared to the isolated cells, which lyse and die out sooner. The authors show that when L-forms exist as isolated single cells, the growth in volume exceeds the rates at which surface area increases, leading to lysis. The authors further quantify the circularity and effective radius in growing proto-colonies, qualitatively estimate membrane tension and suggest that the confined space allows for mechanical shear in these cells. They propose that the mechanical stress on the membranes from adjacent cells in confined spaces deforms membranes and supports cell division to keep the population growing. These findings are also supported by modelling the proto-colonies in quasi-2D planes.

      The study is quite interesting and significant as it has implications for both evolutionary aspects as well as clinical importance, given the proliferation of certain pathogens as L-forms. The aspect of carrying out long-term imaging of colonies of L-forms as spatially constrained entities and the findings are fascinating. While the conclusions presented are backed by experiments, I only have a few questions concerning the proposed mechanism of division and proliferation of these proto-colonies.

      (1) The authors propose that the growth of neighbours leads to shearing forces in membranes and show that membrane tension increases at the periphery of the proto-colonies. They suggest that the increased membrane tension leads to a greater chance of deformation, enabling cell division. However, it is not quite clear how greater membrane tension could lead to cell division. Studies have suggested that membrane fluidisation is important for the cytokinesis event, which includes FtsZ-based division (Ramirez-Diaz, 2025).

      (2) Thus, it becomes quite important to rule out any role for the cytoskeletal proteins in the observed division with an increase in membrane tension. The authors note in line 188 that the division in protocells is independent of FtsZ, but this independence is for protocells that divide by extrusions and resolution, where the membrane is highly fluidised (Mercier et al., 2012).

      (3) The authors may use the L-form derivative where the FtsZ protein can be depleted and assess the proliferation of the proto-colonies. Likewise, authors should rule out the role of MreB as well.

      (4) Although the growth rates have been shown to be similar for proto-cells and the proto-colonies, and only the membrane tension has been shown to be higher at the periphery, it is also important that the authors rule out any increased lipid synthesis in the fraction of dividing cells in these proto-colonies. Without this, one could also envisage a model where membranes are fluidised due to an increase in lipid biosynthesis in a fraction of cells in these confined spaces, leading to increased vesiculations which experience membrane shear and deform. The authors can also consider examining proto-colonies of L-forms of branched-chain fatty acid-deficient strains.

      (5) Lastly, why does CellROX stain the proto-colonies? Are these tightly packed cells experiencing higher oxidative stress, and could that also contribute to membrane tension? This should at least be discussed.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript "Mechanical interaction enables a collective mode of protocell proliferation" addresses an interesting and potentially high-impact question about protocell proliferation in prebiotic environments. The central observation that wall-deficient B. Subtilis proliferate in dense colonies but die by membrane rupture in isolation is striking and a fundamental contribution to the field. However, the data and the mechanistic explanation offered for this observation are incomplete. The measurement and analyses used to build the mechanistic case raise methodological questions that may be difficult to fully resolve with the existing data and approach, and the authors should therefore consider whether additional independent experiments are needed to support the mechanical shearing hypothesis.

      Strengths:

      The central observation that wall-deficient B. Subtilis proliferate in dense colonies but die by membrane rupture in isolation is convincing and a significant contribution to the field interested in the growth of protocells. This adds an important aspect of collective growth that is different from individual dynamics.

      Weaknesses:

      (1) The surface-volume balance ratio η is an elegant concept and provides an intuitively reasonable framework for understanding why isolated cells lyse. However, its application here rests on treating cells as flat discs of uniform thickness, and Figure S4 makes clear that the cells are highly irregular and lobulated in ways that make this approximation questionable. The authors should clarify whether they have validated this assumption, for instance, through direct thickness measurements or sensitivity analysis. However, even with such validation, the modest quantitative differences between aggregated and isolated η trajectories, combined with the inherent difficulty of accurate perimeter measurement in these morphologically complex cells, mean that η measurements are unlikely to provide robust quantitative support for the mechanism. The authors should therefore consider whether η is better presented as a motivating conceptual framework rather than primary quantitative evidence and seek more direct experimental support for the surface-volume balance argument through independent means. For instance, osmotic pressure manipulation to test whether reducing volume expansion pressure preferentially rescues isolated cells.

      (2) The comparison of circularity between colony and isolated cells is complicated by the fact that the segmentation approach is fundamentally different in the two conditions; isolated cell boundaries are detected against a clear background, while colony boundaries are detected from inter-cell fluorescence gradients. The authors should address whether this introduces systematic bias. However, this may be difficult to fully resolve given the inherent complexity of the system, and that the deformation-division correlation in Figure 3C, while suggestive, would be substantially strengthened by a more direct perturbative approach. Specifically, can cell deformation be mechanically induced in isolated cells, for instance, using micromanipulation, external flow, or confinement in fabricated microstructures, to test whether artificially deformed isolated cells gain the ability to divide? Such an experiment would provide direct evidence for the deformation-division link that the correlational analysis cannot.

      (3) The interpretation of FliptR lifetime as a direct membrane tension readout is complicated in this system because cell-cell interfaces contain two apposed bilayers in proximity, potentially altering FliptR photophysics through changes in local membrane density and dielectric environment independently of tension. The authors should address whether they have considered this possibility and what controls were performed. Disambiguating tension-dependent from environment-dependent lifetime changes is technically challenging and suggests that the membrane tension argument would be more convincingly supported by an independent measurement approach. For instance, tether-pulling experiments using optical tweezers on isolated versus colony cell membranes, or testing whether membrane tension-modulating interventions such as osmotic shifts produce the predicted changes in cell fate, would provide more direct evidence. The current FLIM data should be regarded as suggestive rather than conclusive.

      (4) The Cellular Potts Model reproduces the experimental observations, but since its key parameters, particularly the substrate-pinning energy, were calibrated against those same observations, this demonstrates internal consistency rather than independent validation. The η-based lysis criterion is implemented as a model input, meaning the model cannot independently confirm the η hypothesis. The authors should clarify the extent to which model parameters were fitted to data versus independently motivated and be explicit that the model is best understood as a mechanistic illustration rather than independent evidence.

    4. Reviewer #3 (Public review):

      Summary

      This manuscript reports that protocells derived from wall-deficient B. subtilis proliferate well when densely packed but fail to divide and eventually lyse when isolated. The authors attribute this density-dependent proliferation to mechanical shearing between growing neighbors, which deforms cells and increases the likelihood of membrane stalk formation and subsequent scission, enabling division without any dedicated molecular machinery. Through a combination of quantitative imaging, membrane tension measurements, and Cellular Potts Model simulations, the authors make a compelling case that self-generated mechanical stresses are critical for sustaining population growth in protocolonies. The findings have implications for understanding the lifestyles of primitive life forms, L-form bacterial pathogenesis, and the design of synthetic cells.

      Strengths

      The central finding is both surprising and counterintuitive: crowding is not just tolerated by protocells but is required for sustained population growth. The mechanism the authors propose is interesting: mechanical shearing between growing neighbors deforms cells, increasing the likelihood of membrane stalk formation and thus division, all without dedicated molecular machinery. Conceptually, this is a type of biophysical "scaffold" (Jacobeen et al. 2018, Nat. Phys.; Day et al. 2022, Biophys. Rev.) in which key elements of a Darwinian loop, namely a life cycle involving growth and reproduction, are provided "for free" by physics, enabling open-ended Darwinian evolution that can eventually bring these life cycle components under developmental control. Such scaffolds, both biophysical and ecological (Black et al. 2020, Nat. Ecol. Evol.; Libby & Rainey 2013, Phys. Biol.), are likely key mechanisms in the origin of life and in evolutionary transitions in individuality, and this paper provides a nice example of how they can work in a protocell context.

      The combination of experiments and modeling works well. The membrane tension measurements are the strongest piece of evidence for the proposed mechanism, showing directly that tension is elevated in protocolonies and concentrated at cell-cell interfaces. The Cellular Potts Model captures the key experimental features. The discussion is nicely balanced, particularly the note about Gram-negative L-forms, whose rigid outer membrane may preclude this mechanism, which is a testable prediction for future work. I would suggest the authors also discuss the connection to biophysical scaffolding, as I think this is conceptually important and would help situate their work within a broader framework for understanding how primitive life cycles can arise from physical processes (see also Zamani-Dahaj et al. 2023, Genes; Hammerschmidt et al. 2014, Nature).

      Weaknesses

      The surface-volume balance analysis is central to the argument, and it depends on the assumption that cells have a fixed thickness of 0.8 µm, taken from the width of walled cells. But these are wall-deficient cells, which are mechanically quite different, and their thickness could plausibly vary during growth or under compression. I think the paper would benefit from either a direct measurement of cell thickness or a sensitivity analysis showing how η responds to plausible variation in this parameter. If the results are robust, that would put the analysis on much firmer ground.

      The positive correlation between cell shape deformation and division rate (Figure 3C) is central to the proposed mechanism, but I think the paper needs to be more careful about the jump from correlation to causation. The authors propose that deformation increases the likelihood of membrane stalk formation, leading to scission. That is plausible, but an alternative is that cells with higher local growth rates both deform more and divide more frequently, with the two outcomes driven independently by the same underlying cause. The paper does show that average volume growth rates are indistinguishable between aggregated and isolated cells, which argues against a simple "faster growth explains everything" interpretation, but this does not rule out local variation within protocolonies driving the correlation. I think the most convincing experiment would be to apply external mechanical stress to isolated cells and see if that alone can drive division, decoupling deformation from growth. I realize that this may be technically very difficult, but at a minimum, the paper should acknowledge this as an alternative hypothesis.

      The Cellular Potts Model has quite a few free parameters (Table S1), and it is not clear how tightly these are constrained by the data. A sensitivity analysis would go a long way toward showing that the results are robust and not overly dependent on specific parameter choices.

      In any case, this is a strong paper with a cool finding and an interesting mechanistic explanation. I think it will be of broad interest, particularly to people thinking about the origins of life and synthetic cell design.

    1. eLife Assessment

      This important study combines experimental evidence with computational modeling to suggest that neurons in the frontal eye field form two overlapping topographical maps, one for visual inputs and the other for eye movement directions. The experiments leverage the smooth cortex of the marmoset to provide solid evidence that two maps exist at different scales. The modelling work provides a potential mechanistic insight into how these patterns may be used for flexible visuomotor transformations.

    2. Reviewer #1 (Public review):

      Summary:

      Flexible natural behavior requires flexible sensory-motor mapping. In the visual domain, a visual stimulus at one location can guide a saccade toward another. How the receptive field (RF) and motor field (MF) properties of oculomotor structures support this flexibility is not known. Dotson and Reynolds address this question in the marmoset, using oblique Neuropixels penetrations across horizontal segments of the frontal eye field+, supplemented by electrical microstimulation. They report that visual RF and saccade MF vector angles each change smoothly with occasional abrupt jumps, that the two maps are organized as mosaics at distinct preferred spatial scales, and that a moiré interference pattern arising from a constrained spatial-scale mismatch between partially correlated mosaics reproduces the empirical distribution of RF-MF angular differences. They conclude that visuomotor flexibility is embedded in the geometry of mismatch and matches between visual and motor maps.

      Strengths:

      (1) The question is well-motivated. Sensory-motor mapping is known to be flexible, and asking whether the topographic relationship between the two maps itself supports that dissociation is a fresh reframing of a long-standing problem in oculomotor control.

      (2) FEF+ lies on the smooth marmoset cortical surface, which permits high-density horizontal sampling that would be difficult in the macaque arcuate sulcus, and oblique penetrations are a sensible way to track tuning across the surface. The dataset is substantial by the standards of the field (39 sites of high-density recordings across two animals, several thousand isolated units).

      (3) The data are thorough, and the convergence of three independent lines of evidence is the strongest feature of the paper. Unit recordings, electrical microstimulation, and two architecturally distinct generative models point to the same organization.

      (4) The central idea is conceptually novel. The proposal that flexibility can reside in the geometry of the maps, rather than only in time-varying activity, is original, and it generates concrete, testable predictions for tasks that require flexible visuomotor routing.

      Weaknesses:

      Major concerns

      (1) The analysis collapses each oblique penetration onto a single horizontal axis and pools angles across all cortical layers, treating cortical distance as purely tangential. Because the trajectory is angled, horizontal distance and depth are confounded, so some of the apparent RF-MF drift along a penetration could reflect a laminar transition, in addition to tangential mosaic structure.

      (2) RFs and MFs are estimated from the same free-viewing sessions in temporally adjacent epochs, leaving each measurement open to contamination by the other. Activity near a saccade can reflect peri-saccadic remapping rather than the stable retinotopic RF, and saccade-aligned activity following a recent flash can carry a residual visual component, given the long-lasting visual responses in FEF+ (>500 ms). Residual cross-contamination of this kind would tend to make RF and MF angles look more similar than they are, inflating the apparent local coupling and biasing the RF-MF difference distribution that the moiré model is fit to.

      (3) The paper claims that visual and saccade mosaics occupy distinct spatial scales, but the two preferred spatial frequencies are close, and the separation is summarized by overlapping "failed-test" bands rather than by a statistical test or confidence interval on the preferred frequency itself. The reliability of this separation is not established.

      (4) It is not clear whether the moiré model is a better model than the non-mosaic alternative. The moiré models are shown to be consistent with the data through failure to reject a Kolmogorov-Smirnov null, which is a weak form of evidence, and they are not benchmarked against a non-mosaic alternative or null model. The AM/NM convergence demonstrates architecture independence, but not that a mosaic organization is required.

      Minor concerns

      (1) The link from topography to behavioral flexibility (such as anti-saccades and other context-dependent transformations) is presented as a prediction but is not tested with any task manipulation. The work establishes an organizational principle and a plausible generative mechanism; whether that organization is actually exploited during flexible behavior remains open, and the framing should make this clear so the functional claim is not over-read.

      (2) It is unclear how relative depth (depth 0) is defined and how layer boundaries were assigned. The Methods mention common-average re-referencing for CSD and local field analyses, but no CSD or power-depth profile is shown to anchor the layer IV / depth 0 reference across penetrations.

      (3) The Discussion is brief relative to the strength of the claims. It would be helpful to address the concerns and alternative explanations above, where these cannot be fully resolved by the data.

    3. Reviewer #2 (Public review):

      Summary:

      The authors asked how the visuomotor system can keep visual selection and saccade targeting related but not rigidly coupled-the flexibility required for tasks like anti-saccades. They recorded visual receptive fields (RFs) and saccade motor fields (MFs) from individual neurons in marmoset frontal eye fields and adjacent premotor eye fields (FEF+) using obliquely inserted Neuropixels probes that traverse horizontal segments of the smooth marmoset cortex. They reported that visual and motor vector angles each changed smoothly with occasional abrupt jumps (a mosaic, rather than retinotopic, organization), that the two maps drifted with respect to one another, and that they were best described by mosaic-map models tuned to different preferred spatial frequencies. They then proposed that the offset in spatial scale, combined with partial shared structure between the maps, produced a moiré interference pattern, in which the distribution of local visual-motor angle differences matched the data.

      Strengths:

      (1) Unlike in the macaque brain, where FEF is buried in the arcuate sulcus, the marmoset cerebral cortex is lissencephalic (smooth). The authors used this feature to their great advantage and sampled horizontal mesoscale structure with oblique penetrations of ultra-high-density Neuropixels probes.

      (2) The mosaic framing was grounded in previous studies on direction maps in ferret V1 and MT, and the rate-of-change analysis (Supplementary Figure 2) plausibly reproduced the fracture-line phenomenon of those maps.

      (3) The authors confirmed the robustness of the finding by reaching the same conclusions using two architecturally distinct generators: the Fourier-based annulus model and the Gaussian-noise model.

      (4) They additionally provided an independent confirmation of the mosaic saccade-vector organization using electrical microstimulation. They reconciled the lower microstimulation spatial frequency by matching the spatial-averaging footprint (Supplementary Figure 6).

      (5) The Noise Mosaic (NM) model was used thoughtfully to decouple two properties that the Annulus Mosaic (AM) model confounds - spatial-scale offset and inter-map correlation - and to show that both an intermediate correlation (ρ ≈ 0.6) and a scale offset are required. Conceptually, a structural (topographic) substrate for visuomotor flexibility is a fresh alternative to the standard account in which flexibility lives entirely in time-varying activity on a single map.

      Weaknesses:

      The two claims here are not of the same strength. The first claim that RF and MF angles are organized as mosaics at distinct spatial scales was well supported. The second claim that a moiré interference pattern is the substrate for visuomotor flexibility was an inference rather than a direct observation, and several features of the design contribute to how strongly it can be held.

      First, the recordings were one-dimensional. Each oblique penetration yielded a line through the cortex, so the two-dimensional moiré pattern (Figure 4A) existed only in simulations; it was not reconstructed from the data. What the data provided was the marginal distribution of local angular differences, and the model was accepted when its simulated distribution was statistically indistinguishable from the empirical one. Matching a low-dimensional summary statistic is necessary but not strongly sufficient - multiple underlying architectures could produce similar 1-D difference distributions - so the moiré interpretation is best read as a plausible and parsimonious interpretation of the data rather than a confirmed mechanism.

      Second, the inferential logic needs to be strengthened. The "preferred" spatial frequencies were those at which a two-sample Kolmogorov-Smirnov test fails to reject equality between model and data. Failure to reject is not confirmation, and the width of the accepted band depends on statistical power, which depends on sample size. The authors did show that most parameter combinations were rejected, so the test did discriminate. That said, a continuous goodness-of-fit landscape with confidence intervals on the preferred SF, and a direct test that the visual and saccade preferred SFs differ, would better support the "distinct spatial scales" claim than visual inspection of two overlapping troughs.

      Third, the dataset was from two male marmosets, with 18 of 39 sites contributing to the core analyses, and the angular-difference distributions were pooled across penetrations and animals. This is standard for primate electrophysiology, but it means the spatial statistics were assumed stationary across the region, and individual variability in map layout was averaged over. The oblique-penetration geometry also added some uncertainty: the spatial-frequency estimates (in cycles/mm) are only as accurate as the reconstructed penetration angles (18.2{degree sign} {plus minus} 8.3{degree sign}), and angle error would propagate directly into the inferred scales.

      Fourth, while the authors suggested the moiré interference pattern can serve flexible routing for behaviours such as anti-saccades, this was not directly tested. Instead, they used free viewing and natural saccades, so the paper demonstrated a candidate substrate without testing whether behaviour employs it. This does not undercut the main findings, but readers should treat the functional narrative in the introduction and discussion as a set of predictions rather than results.

    1. eLife Assessment

      In this valuable study, Bhojappa et al. investigate the function of four septin-associated kinases (Elm1p, Gin4p, Hsl1p and Kcc4p) in controlling septin ring stability, actomyosin ring organization and constriction, and cytokinesis in budding yeast. The evidence supporting the conclusions is convincing. The results are a meaningful addition to previously published work and provide a clearer picture of a complex mechanism involving proteins of partially redundant functions.

    2. Reviewer #1 (Public review):

      Summary:

      The authors wanted to better understand how the various septin-associated kinases contribute to septin organization and function in budding yeast. This question has been recently addressed by similar kinds of studies but there are still some open questions, particularly as regards to what extent the kinases may interact with and/or modify components of the contractile ring that drives cytokinesis.

      Strengths:

      This study uses sensitive imaging with good temporal and spatial resolution to monitor the localization of various proteins in living cells. Particularly informative is the use of a GFP/GFP-binding-protein "tethering" approach to ask if the requirement for one protein can be bypassed by physically tethering another protein to a third protein. Results from a yeast two-hybrid assay for measuring protein-protein interactions in vivo are buttressed by direct in vitro binding assays using purified proteins, which is important given the likelihood of "bridging" interactions between yeast proteins in the two-hybrid approach. The authors' conclusions are quite well supported by the data.

      Weakness:

      Ultimately, while the study provides some interesting and novel insights, we still don't understand which phosphorylation events on which proteins are important for the events occurring at the molecular level, so the advance in knowledge is somewhat incremental.

    3. Reviewer #2 (Public review):

      Summary and strengths:

      In this study, Bhojappa et al. investigate the roles of the septin-associated kinases Elm1, Gin4, Hsl1, and Kcc4 in septin organization and cytokinesis in budding yeast. Through quantitative analyses of kinase localization dynamics, septin organization, actomyosin ring (AMR) constriction, and cell morphology, the authors demonstrate that Elm1 and Gin4 play particularly important roles in maintaining proper septin architecture and cytokinetic progression. The work further identifies an interaction between the Gin4 KA1 domain and the Hof1 F-BAR domain and provides evidence that several cytokinesis-related functions of Gin4 are independent of its kinase activity. Artificial tethering approaches further supports that spatial organization at the bud neck is critical for the execution of septin-dependent cytokinetic processes. The authors combine live-cell imaging, quantitative analyses, biochemical interaction assays, and genetic perturbations to build a comprehensive framework for understanding how these kinases contribute to cytokinesis.

      Comments on revised version:

      The revised manuscript has been substantially improved. The authors have carefully addressed the concerns raised during review by providing additional experiments, analyses, quantifications, clarifications, and improved presentation of the data. The conclusions are now well supported by the experimental evidence. The study advances our understanding of the mechanisms linking septin organization to cytokinesis and will be of interest to researchers studying septins, cell division, and cytoskeletal regulation.

      Overall Assessment:

      This work provides valuable mechanistic insight into the coordination of septin organization and cytokinesis by septin-associated kinases. The experiments are carefully executed, the analyses are thorough, and the conclusions are supported by the data presented. The manuscript represents a useful contribution to the fields of cytokinesis and septin biology.

    4. Reviewer #3 (Public review):

      Summary:

      The study by Bhojappa et al. brings new and interesting elements about the stability of the septin ring and the crosstalk between septin and actomyosin ring assemblies. The study focuses on the four kinases associated with the septin ring, Elm1p, Gin4p, Hsl1p and Kcc4p. Elm1 and Gin4 show the strong knock-out phenotypes, whereas Hsl1p and Kcc4p show the weak knock-out phenotypes. The Elm1p/Kccp1p and Gin4p/Hsl1p pairs show similar timing at the bud neck. While these kinases share redundant functions, Gin4 appears to have a unique interaction with the BAR domain protein Hof1, revealing a novel direct interaction between the septin and actomyosin rings. Interestingly, the kinase activity of Gin4 is not required for its role in septin organisation and AMR constriction. The last part of the manuscript shows an original protein tethering protocol used to show that Hsl1 and its membrane binding ability are required for phenotype rescue of gin4null cells.

      Comments on revised version:

      I thank the authors for their thorough and thoughtful response to my review. The revised manuscript clearly reflects their efforts to provide rigorous and high-quality science. Addressing the concerns raised in the review required significant effort, but the improvements in the manuscript make it clear that the work was well worth it.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors wanted to better understand how the various septin-associated kinases contribute to septin organization and function in budding yeast. This question has been recently addressed by similar kinds of studies but there are still some open questions, particularly as regards to what extent the kinases may interact with and/or modify components of the contractile ring that drives cytokinesis.

      Strengths:

      This study uses sensitive imaging with good temporal and spatial resolution to monitor the localization of various proteins in living cells. Particularly informative is the use of a GFP/GFP-binding-protein "tethering" approach to ask if the requirement for one protein can be bypassed by physically tethering another protein to a third protein. Results from a yeast two-hybrid assay for measuring protein-protein interactions in vivo are buttressed by direct in vitro binding assays using purified proteins, which is important given the likelihood of "bridging" interactions between yeast proteins in the two-hybrid approach. The authors' conclusions are quite well supported by the data.

      Weaknesses:

      A control for non-specific binding is missing from the in vitro binding assay. The figures suffer sometimes from the very small text in the labels, which obscures understanding. Ultimately, while the study provides some interesting and novel insights, we still don't understand which phosphorylation events on which proteins are important for the events occurring at the molecular level, so the advance in knowledge is somewhat incremental.

      We thank the reviewer for highlighting the strengths of our imaging pipelines and protein-protein interaction data. We have now included appropriate controls for the in vitro binding assays, which demonstrate that the observed interactions are specific (Fig. 2H). We have also revised all figures to improve clarity, including increasing font sizes to enhance visibility across panels. We agree that mapping the specific phosphorylation sites regulated by these septin kinases would provide valuable mechanistic insights. However, only a few studies have addressed this direction so far (Mortensen et al., 2002; Asano et al., 2006 and Marquardt et al., 2024) [1-3]. The current study focuses on the interplay among septin-associated kinases and their role in regulation of the actomyosin machinery (AMR). In this context, we highlight several key findings:

      (i) a molecular link between the septin kinase network and AMR through physical interaction between the KA1 domain of Gin4 and F-BAR domain of Hof1 (Fig. 2F-2H and S3H);

      (ii) a kinase-independent role for Gin4 in coordinating septin organization and AMR dynamics (Fig. 3A-3E, S3F & S3G, S3I and 4F-4H);

      (iii) a novel role for Hsl1 in regulating septins and the AMR downstream of Gin4 and Elm1, potentially through plasma-membrane binding (Fig. 4I-4K, 5A-5F, 8A-8G and S9H-S9J); and

      (iv) crosstalk between Gin4 and Hsl1 that is independent of their role in the morphogenetic checkpoint (Fig. 6A-6D and S6A-S6C).

      We have clarified this scope in the Discussion section and explicitly stated that mapping these phosphorylation sites will be an important direction for future work (Lines 623-626).

      Reviewer #2 (Public review):

      Summary:

      In this paper, Bhojappa et al. provide insights into the function of septin-related kinases Elm1, Gin4, Hsl1, and Kcc4 in septin organization and actomyosin ring (AMR) structure and constriction. Their findings are both corroborative of and complementary to previous related studies.

      First, the authors provide a comparative analysis of the dynamic localization of these kinases at the bud neck, as well as a comparative analysis of defects in septin localization, splitting dynamics, AMR constriction rates, and cell morphology in kinase deficient cells. They find that septin localization and splitting kinetics, as well as AMR constriction rates, are significantly perturbed in elm1∆ and gin4∆ mutants but remain largely unaffected in hsl1∆ and kcc4∆. A similar trend is observed in terms of cell morphology and viability.

      Next, the authors focus on elm1∆ and gin4∆ cells, demonstrating that the residence time of the F-BAR protein Hof1 is significantly increased and defective in these mutants. Using yeast two-hybrid (Y2H) and in vitro binding assays, they show that the KA1 domain of Gin4 interacts with the F-BAR domain of Hof1, which may explain the cytokinesis-related functions of Elm1 and Gin4. Supporting this, they find that Gin4's role in septin localization, AMR constriction kinetics, and Hof1 bud neck localization is kinase-independent.

      The authors then conduct a series of artificial tethering experiments given their bud neck localization is mostly interdependent. They first demonstrate that artificially tethering Gin4 to the bud neck rescues the morphology defects of elm1∆ cells, with the strongest rescue observed when Gin4 was forced to interact with Hsl1-an effect that was also kinase-independent. Additionally, artificial tethering of Hsl1 to the bud neck restores the morphology of elm1∆ cells in a KA1 domain-dependent manner, suggesting that Hsl1 functions downstream of Elm1 to maintain normal cell morphology. Consistently, artificial tethering of Elm1 to the bud neck in gin4∆ cells rescues morphology defects, as well as defects in Myo1 localization and AMR constriction, but only in the presence of full-length Hsl1. The rescue fails in the absence of Hsl1 or when using a version of Hsl1 lacking the KA1 domain, which supports the role of Hsl1 downstream to Elm1 in cytokinesis.

      Strengths:

      Altogether, this study offers valuable insights into the mode of cytokinesis regulation mediated by the septin-related kinases, mainly Elm1, Gin4, and Hsl1, and would be an important contribution to the field of septins and cytokinesis after addressing current weaknesses.

      We thank the reviewer for the detailed summary and for highlighting the novel findings of our study.

      Weaknesses:

      (1) When assessing rescue of the elm1∆ phenotype, it needs to become clearer whether only morphology or also cytokinesis and septin organization are rescued.

      To clarify the extent of rescue observed in elm1Δ cells, we extended our analysis beyond morphological parameters by quantifying septin organization and AMR constriction dynamics. These analyses now show that artificial tethering partially restores septin organization and AMR constriction kinetics in elm1Δ cells in addition to improving cell morphology. These results are now described in detail in Fig. 5 and lines 373-406.

      (2) The quantification of the microscopy data does not always match up with the example images, and it's not always clear how the authors quantitatively analyzed their data.

      We revised the manuscript to clearly outline the quantification methods used for microscopy data analysis and specified the statistical tests, number of cells analyzed, and number of experimental replicates in the figure legends. We also clarified the criteria used for phenotype scoring and quantification in the Materials and Methods section. In addition, We replaced representative images where necessary to accurately reflect the quantified data throughout the revised manuscript.

      (3) The forced tethering data are key to the paper, but the lack of a summarizing table makes it difficult to grasp the full picture.

      We agree with the reviewer and have now included a new summary table (Table 1) that compiles the results of all artificial tethering experiments presented in this study, including the percentage of rescue in cellular morphology observed upon forced tethering of these kinases to the bud neck, thereby providing a clearer overview of these experiments.

      (4) Novel results and those confirming earlier results could be better distinguished.

      We have improved the overall clarity of the manuscript to distinguish novel findings from the results that corroborate previous studies, and have cited the appropriate literature throughout the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      The study by Bhojappa et al. brings new and interesting elements about the stability of the septin ring and the crosstalk between septin and actomyosin ring assemblies. The study focuses on the four kinases associated with the septin ring, Elm1p, Gin4p, Hsl1p, and Kcc4p. Elm1 and Gin4 show strong knock-out phenotypes, whereas Hsl1p and Kcc4p show weak knock-out phenotypes. The Elm1p/Kccp1p and Gin4p/Hsl1p pairs show similar timing at the bud neck. While these kinases share redundant functions, Gin4 appears to have a unique interaction with the BAR domain protein Hof1, revealing a novel direct interaction between the septin and actomyosin rings. Interestingly, the kinase activity of Gin4 is not required for its role in septin organisation and AMR constriction. The last part of the manuscript shows an original protein tethering protocol used to show that Hsl1 and its membrane binding ability are required for phenotype rescue of gin4null cells.

      Strengths:

      The combination of genetics, cell imaging, and biochemical characterization of proteinprotein interactions is attractive.

      We thank the reviewer for recognizing the significance of our findings and for the helpful suggestions.

      Weaknesses:

      (1) Imaging and data analysis is the main weakness of this manuscript. The authors must avoid manual counting and selection when easy analysis software can be used to limit bias. Instead of presenting unclear statistics of "percentage phenotypes", they need to define clear metrics to offer meaningful phenotype analysis.

      We agree that improving the quantitative rigour of the image analysis is essential for this study. Accordingly, we implemented a semi-automated image analysis workflow in the revised manuscript that defines reproducible metrics, such as aspect ratio, and reduces reliance on subjective phenotypic scoring. The inclusion of these parametric measurements enables clearer and more objective comparison of the rescued phenotypes.

      (2) This manuscript examines a very complex mechanism with four kinases of overlapping function using new data and existing literature. A clearer picture/model at the end of the manuscript that synthesizes the current knowledge would be beneficial:

      We incorporated a new representative model (Fig. 9) that integrates current knowledge in the field with our findings. This model highlights crosstalk among Elm1, Gin4, and Hsl1 as a key mechanism coordinating septin architectural transitions with AMR constriction during cytokinesis and is discussed in lines 520-536 of the revised manuscript.

      We sincerely thank all the reviewers for their insightful comments. We incorporated new results in Fig. S3A, S3B, S4A-S4F, 5A-5F, 6A-6D, S6A-S6C, 8E, 8F, S10B and 9, along with Table 1 summarizing the artificial tethering experiments in the revised manuscript. We believe that these revisions have improved the rigor of our analyses and enhanced the overall clarity of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 70: " abnormal cytokinesis defects": this language is redundant, either "abnormal" or "defects" would suffice:

      We thank the reviewer for identifying the redundant wording. We have rephrased the final paragraph of the Introduction section.

      Please refer to line numbers 86-97.

      (2) Currently, the final paragraph of the Introduction is an extensive, detailed summary of the results. This is unnecessary, as the Abstract and the Results sections summarize the results. Better would be a short statement of the questions addressed in the manuscript:

      We thank the reviewer for this suggestion. We have revised the final paragraph of the Introduction to remove the detailed summary of results and instead outline the key questions addressed in this study. The revised paragraph now emphasizes the knowledge gap regarding how septin-associated kinases regulate septin organization and coordinate cytokinesis. Please refer to line numbers 86-97.

      (3) In the Results, the wording of the heading associated with the first section is confusing. "and their defects during the cell cycle": it is not expected that the kinases themselves will have defects; defects may be observed in cells upon mutation of the kinases, but that is not clear from this wording:

      We thank the reviewer for this comment. The heading has been revised from “and their defects during the cell cycle” to “and defects associated with their deletions” for improved clarity.

      Please refer to line numbers 99-100.

      (4) Lines 107-109 (" they may act as molecular signals for the transition and a trigger for crosstalk of cytokinesis") are redundant with earlier lines 104-105 (" suggesting them to be a possible trigger for septin remodelling"):

      We have rephrased the text to improve clarity and avoid redundancy.

      Please refer to line number 123-125.

      (5) Lines 121-122: "Septin-associated kinases are believed to play an essential regulatory role": is the role essential, or is it regulatory? Since these kinases are not individually essential for cytokinesis, "essential" doesn't seem appropriate here:

      This has been corrected in the revised manuscript.

      (6) Figure 1 panels C and F: the font size is exceedingly small for the labels and should be greatly increased. This is true for multiple panels in Figures 2-5 and the Supplemental Figures as well:

      We thank the reviewer for highlighting this issue. We have significantly increased the font sizes of the text and the x- and y-axis labels across all figures in the manuscript.

      (7) The timing of mitotic spindle breakdown was used as a timepoint for comparison but it is not made clear in the manuscript how this was determined. Presumably, the mRuby2-Tub1 marker was visualized and the timepoint when the mitotic spindle separated into two discrete entities was called the "breakpoint" timepoint, but it would be important to better describe (and ideally show an example) of how that timepoint was determined. The only images I find with labelled tubulin are Tub1-GFP and these do not show spindle breakdown:

      We thank the reviewer for raising this point. We used GFP-Tub1 (pAFS125-GFPTUB1) and mRuby2-Tub1 (pHIS3p:mRuby2-Tub1+3′UTR::URA3) plasmids to visualize spindle dynamics across the cell cycle, and in both cases defined the spindle breakpoint as time zero. As suggested, we have now included time-lapse images of Cdc3-mCherry and GFP-Tub1 in wild-type cells in Fig. S2D to illustrate the spindle breakpoint event used for temporal alignment. Corresponding changes have also been made in the Materials and Methods section to explicitly describe this analysis.

      Please refer to line numbers 760-761.

      (8) Line 165: "F-BAR protein Hof1, which senses and induces membrane curvature" and lines 170-172 "the F-BAR protein Hof1, which is known to be associated with septin hourglass and transit to AMR during split ring trigger". It is awkward to introduce the same protein twice, in two different ways, within a few lines of each:

      We thank the reviewer for noting the redundant description. We have rephrased the text to introduce the F-BAR protein Hof1 in a more concise and streamlined manner while retaining the relevant functional information.

      Please refer to line numbers 206-208.

      (9) Throughout, it would be helpful to introduce more line breaks and organize the Results into smaller paragraphs:

      We thank the reviewer for this suggestion. We have revised the layout of the Results section and introduced additional line breaks to improve readability.

      (10) This is somewhat of a personal preference, but in the interest of transparency (and with the understanding that the 0.05 value is entirely arbitrary), would authors be willing to show actual P values in the figure panels rather than "**" or "ns", for example? Some readers may wish to apply a different standard of significance than 0.05, and not showing the P values makes this impossible. Furthermore, some readers (like this one) may interpret differently a P value of 0.044 versus "*" or 0.051 vs "ns":

      We agree with the reviewer’s suggestion. To improve the transparency of the quantitative analyses, we have modified the graphs across all main and supplementary figures to display the “actual p-values” and specified the corresponding statistical tests in the figure legends, with significance indicated by asterisks. An example is provided in the attached image showing the residence time of Inn1-mNG in wild-type and gin4Δ cells complemented with kinase-active and kinase-dead Gin4 constructs (Fig. 3E in the revised manuscript). For this analysis, significance was assessed using the nonparametric Kruskal-Wallis statistical test.

      (11) It is written that the authors "performed a Yeast Two Hybrid screen to find novel interacting partners at the bud neck" but I do not find evidence anywhere of a "screen" being performed, i.e., an unbiased search of many proteins to find a few interactors. Instead, it appears that the authors performed a two-hybrid assay to visualize interactions between a small, specific set of proteins. "Screen" should be replaced with "assay", as is currently the case in the Methods section. It is probably also valuable to point out in the text that using a yeast two-hybrid assay to assess interactions between yeast proteins has the caveat that any interactions observed could be indirect, as they may be "bridged" by endogenous yeast proteins.

      We thank the reviewer for raising this point. Our Yeast Two-Hybrid experiments were performed using a small subset of septin-associated proteins rather than as an unbiased screen. Accordingly, we have replaced the term “screen” with “assay” throughout the revised manuscript and in the Materials and Methods section.

      Please refer to line 241.

      We also agree with the limitations inherent to this assay and have now explicitly stated it in the revised manuscript, please see lines 248-255.

      (12) Panel 2H: I do not understand what the middle (as opposed to the top and bottom) blot segment represents. It is labelled "anti-HIS", like the one above it, but it is not associated with any molecular weight/ladder marker and I do not know what other species in the binding reaction in that lane would be recognized by the anti-HIS antibody. Perhaps the top segment is an "input" sample, and below it (in the middle segment) is what was bound to the beads. The figure legend is uninformative in this regard. Also, panel I in this figure is unnecessary to show, assuming that the bands shown in H are what I think they are. The blot makes the point without the need for quantification.

      As suggested by the reviewer, we have added molecular weight markers for each blot panel showing the input and bead-bound fractions. The figure legend has also been updated accordingly to clearly describe the different blot segments and experimental conditions.

      Please refer to figure legend 2H.

      In addition, as suggested by the reviewer, we have removed the quantification graph corresponding to the in-vitro binding assay from the revised manuscript.

      (13) The results in Figure 2H demonstrate that the Gin4-KA1 fragment is not non-specifically "sticky", because it does not bind GST alone, but there is no demonstration that binding by the Hof1 fragment is specific because there is no equivalent negative control for binding:

      We thank the reviewer for this suggestion. We repeated the in-vitro binding assay using 6His-bdSUMO as a negative control alongside 6His-bdSUMO-Gin4<sup>KA1</sup> to demonstrate binding specificity of the Hof1 fragment. The Hof1 N-terminal F-BAR fragment did not pull down the control 6His-bdSUMO fragment but specifically pulled down 6His-bdSUMO-Gin4<sup>KA1</sup> under identical experimental conditions (Fig. 2H), confirming the specificity of the interaction between Hof1 F-BAR domain and Gin4KA1.

      The corresponding text and results have been updated in the revised manuscript.

      Please refer to Fig. 2H and lines 259-263.

      (14) Line 257-258: "Localisation via Hsl1 is necessary to rescue the morphological defects exhibited by Δelm1 cells partially": what does "partially" refer to here? To the rescue, or the defects?

      The term “partially” refers to the extent of rescue. Specifically, elongated cell morphology was rescued in 63.75% of the elm1Δ cell population, rather than in all cells, upon artificial tethering of Gin4-GFP to the bud neck via Hsl1-GBP.

      (15) Lines 292-293: "can restore the morphological defects": this wording is unclear. "Restore" means "return to a former condition", which in this case would be normal cellular morphology, not defective cellular morphology. Similarly, see line 304: "While Myo1-3xmCherry mislocalisation was restored upon Elm1-GFP tethering": presumably the proper localization was restored, not the mislocalization:

      In Lines 292-293, by phrase “can restore the morphological defects” was intended to indicate rescue of the elongated/clumped morphology associated with gin4Δ cells upon artificial tethering of Elm1-GFP to the bud neck via Shs1-GBP. We have now rephrased this sentence as: “can rescue the elongated/clumped phenotype exhibited by gin4Δ cells”.

      Please refer to line numbers 481-482.

      Similarly, in Line 304, the statement “While Myo1-3xmCherry mislocalisation was restored upon Elm1-GFP tethering” referred to rescue of the Myo1 mislocalization phenotype observed in gin4Δ cells. We have rephrased this sentence as: “However, Myo1-3xmCherry localization was restored to the bud neck upon artificial tethering of Elm1-GFP in gin4Δ cells”.

      Please refer to lines 506-508 in the revised manuscript.

      (16) Lines 295-296: "We find that Elm1-GFP tethering via Shs1-GBP, Bud4-GBP, and Hsl1-GBP" this should be "or", not "and":

      Thank you for pointing this out. We have now corrected the text accordingly.

      Please refer to line number 485.

      (17) Lines 311 and 312 refer to "Inn1-3xmCherry lifetime" but previously Inn1 residence time was measured. Since fluorescence lifetime is a distinct kind of measurement/ assay, it seems important to clarify here what kind of experimental data are being referred to:

      The term “Inn1-3xmCherry lifetime” was intended to describe the residence time of Inn1 at the cell division site, defined as the interval between the initial appearance of the Inn1 fluorescence signal and its complete disappearance during cytokinesis. For clarity and consistency, we have replaced the term “lifetime” with “residence time” in the revised manuscript.

      Please refer to the line numbers 511 and 512.

      (18) Discussion: "Septins are considered as the fourth cytoskeletal elements due to their extensive structural and functional diversity." This sentence is confusing, as it seems to imply that what defines a protein as being "cytoskeletal" is structural and functional diversity rather than anything to do with forming filaments, etc. The rest of this first paragraph of the Discussion also sounds like a summary of the background, is quite redundant with the Introduction, and should be shortened:

      We thank the reviewer for this suggestion. We have rewritten the first paragraph of the Discussion to improve clarity, reduce redundancy with the Introduction, and better emphasize the main findings of the study.

      Please refer to the line numbers 538-552 in the Discussion section.

      (19) Lines 365-366: "Gin4 and Hof1 are synthetic lethal": this should be revised to "gin4∆ and hof1∆ are synthetic lethal". This is a good place to point out that standard yeast nomenclature inserts the ∆ symbol after the gene name, not before (as is done in E. coli genetics, for example):

      We thank the reviewer for this suggestion. We have replaced “Gin4 and Hof1 are synthetic lethal” with “gin4Δ and hof1Δ are synthetic lethal” in the revised manuscript.

      Please refer to line number 576.

      In addition, we have now consistently placed the Δ symbol after the gene throughout the manuscript in accordance with standard yeast nomenclature.

      (20) Line 399: "We also performed an extensive GFP-GBP screens": again, here "screen" implies that a large collection of genes/proteins were assayed, perhaps in an unbiased way, which does not accurately portray what was actually done, which was an extensive tethering study using GFP-GBP:

      We thank the reviewer for this suggestion. We have replaced the term “GFP-GBP tethering screen” with “GFP-GBP tethering assay” throughout the revised manuscript. In addition, we have included a brief description of the specificity and functionality of GBP nanobody and its application in the GFP-GBP tethering strategy, extensively used in this study.

      Please refer to line numbers 326-336.

      Reviewer #2 (Recommendations for the authors):

      Major points:

      (1) Analysis of morphological defects of elm1∆ does not directly reflect the defects in cytokinesis and septin organization. For example, the deletion of SWE1 rescues the morphological defects of elm1∆ cells but not the cytokinesis or septin mislocalization (Bouquin et al., 2000). Considering this, the authors should address whether artificial tethering of Hsl1 to Gin4 in elm1∆ cells rescues the septin and cytokinesis defects or just the morphology. Is the role of Hsl1 in cytokinesis dependent on its role in the morphogenesis checkpoint? Can the authors comment on how much the defects observed by Hsl1 tethering to the bud neck may be a result of bypassing the morphogenesis checkpoint?:

      We thank the reviewer for this important point. We performed time-lapse imaging of Cdc3-mCherry in strains where Gin4-GFP partially rescued the elongated phenotype of elm1Δ cells (63.75%) when tethered to the bud neck via Hsl1-GBP. Under these conditions, 64.29% of cells showed rescue of Cdc3-mCherry mislocalization, and Gin4 localization itself was restored to the bud neck in 57.85% of cells. We also examined Myo1-ymScarletI dynamics, while 76.92% of untethered elm1Δ cells displayed Myo1 mislocalization to the bud cortex, this was reduced to 11.36% upon Gin4-GFP tethering via Hsl1-GBP. Together, these results indicate that the morphological rescue observed in elm1Δ cells is accompanied by restoration of normal septin organization and AMR dynamics.

      Previous work (Bouquin et al., 2000) [4] showed that Swe1 deletion rescues cell elongation in elm1Δ cells without restoring septin organization. Consistent with this, we found that 60.66% of elm1Δ swe1Δ cells exhibited a round morphology, but tethering of Gin4-GFP to the bud neck via Hsl1-GBP in elm1Δ swe1Δ background did not further enhance morphological rescue. These results suggest that the rescue of cellular morphology observed in our tethering experiments may, atleast in part, depend on Hsl1-mediated regulation of the morphogenesis checkpoint.

      Importantly, despite the lack of additional morphological rescue, a clear restoration of septin localization was observed when Gin4-GFP was tethered to the bud neck via Hsl1-GBP in elm1Δ swe1Δ cells. Overall, these results suggest that while Hsl1-dependent morphogenesis checkpoint regulation may contribute to cell shape rescue, the restoration of septin organization is independent of Hsl1’s function in morphogenesis checkpoint and instead reflects a direct requirement for Gin4 and Hsl1 at the bud neck.

      Please refer to Figures 5, 6, and S6 of the revised manuscript for these additional data.

      (2) As suggested by the authors, the interaction of the Gin4-KA1 domain with the FBAR domain of Hof1 may explain the cytokinesis-related functions of Gin4. As an orthogonal approach, how does KA1 domain deletion of Gin4 affect cytokinesis and Hof1 bud neck localization?

      We thank the reviewer for this suggestion. We first examined the bud neck localization of Gin4-ka1Δ-GFP in comparison with full-length Gin4-GFP. We observed that the localization kinetics of Gin4-ka1Δ-GFP were significantly altered relative to the full-length protein, with reduced recruitment and earlier removal from the bud neck. We also analyzed the localization kinetics of Hof1-mNG in both gin4-ka1Δ and gin4Δ cells. Our results show that Hof1-mNG displays increased residence time and altered accumulation kinetics during cytokinesis in both genetic backgrounds. Thus, loss of the KA1 domain phenocopies loss of the full-length Gin4 and is consistent with disruption of the physical interaction between Gin4 and Hof1.

      Please refer to Figure S4 for these results in the revised manuscript.

      (3) The authors state that "Elm1 and Kcc4 were present at lower abundance at the bud neck (Fig S1A-D) compared to the higher abundance of Gin4 and Hsl1, as observed in their fluorescence intensities (Figures S1B-C) ". This is not evident in the figures. The authors should show a quantification of how they judged abundance at the bud neck:

      We thank the reviewer for this question. Quantification of septin kinase fluorescence intensity at the bud neck was performed using the established protocol for measuring protein accumulation kinetics described by Okada et. al. 2020 [5]. Time-lapse imaging for kinetic analysis of septin-associated kinases during bud emergence shown in Fig. S1A-S1D was carried out using a point-scanning confocal microscope with a 100×oilimmersion objective. Different laser intensities were required because the fluorescence signals of Kcc4 and Elm1 were comparatively weak and not readily detectable above cellular background under the imaging conditions used for Gin4 and Hsl1. The images shown in Fig. S1A-S1D are therefore displayed using differential contrast settings to facilitate visualization. We have now explicitly clarified this in the figure legend.

      A more direct comparison of septin kinase abundance at the bud neck is now provided in Fig. S1E-F, where localization kinetics during the HDR transition/septin remodelling stage were captured using a laser-scanning spinning-disk microscope under similar imaging conditions. We have additionally included raw fluorescence intensity profiles during the HDR transition to better illustrate the relative abundance of these kinases at the bud neck during cytokinesis.

      Please refer to Figures S1E and S1F in the revised manuscript for the updated images and quantitative analyses.

      (4) How did the authors determine G1 and M-phase in the experiments shown in Figures S1A-D? Can the authors mark these phases on the timelapse images? Also, How do the authors explain the different behaviour of Cdc3 in graphs S1A-D among different strains?

      We thank the reviewer for this comment. Cell cycle stages were initially inferred based on bud size, where kinase accumulation at the bud neck correspond to bud emergence (small bud, G1), and kinase disappearance coincided with septin splitting (large bud, M phase). However, we agree that accurate assignment of cell cycle stages would require specific cell cycle markers. To avoid confusion, we have removed the cellcycle-specific stage assignments from the Results section and describe the kinetics relative to t=0 (bud emergence).

      The differential dynamics observed in the Cdc3-mCherry kinetic profiles likely reflect heterogeneity within the the cellular population. To address this, we combined the normalized fluorescence intensity profiles of Cdc3-mCherry from the strains expressing GFP-tagged septin kinases during bud emergence and have included this data as reference (Author response image 1).

      Author response image 1.

      Plot showing spatiotemporal kinetics of Cdc3-mCherry in strains expressing either Elm1-GFP, or Gin4-GFP, or Hsl1-GFP, or Kcc4-GFP.

      (5) Forced tethering based experiments are one of the key sets of experiments for this work, but it is difficult to have a comprehensive understanding of all the data considering how large the data set is and how dispersed it is in the supplemental and main figures (Figures 3-4-5 and Figures S4-S5-S6). It would be helpful to provide a table summarizing the tested forced tethering’s and the phenotypic outcome in the tested yeast strains (Wt/mutant):

      We thank the reviewer for recognising the extensive dataset generated from the artificial tethering experiments and for suggesting the inclusion of a summary table. We have now added Table 1, which summarizes the proteins used in the GFP-GBP artificial tethering experiments, their genetic backgrounds, the total number of cells quantified across three independent replicates, and the phenotypic outcomes associated with bud neck tethering under each condition.

      Please refer to Table 1 and lines 342, 361, 451, 453, 456 and 492 in the revised manuscript.

      (6) In the introduction section, it would help the reader to provide more information on already known molecular roles of septin-associated kinases in septin organization and AMR. Later in the results section (i.e. Figure S1, S2, and S6), the authors extensively explain and show data that independently corroborate some earlier findings, which makes it difficult for the reader to distinguish novel findings from the repeated findings. I suggest shortening the text for the corroborative results, which will help to put more emphasis on their novel findings:

      We thank the reviewer for this suggestion. We have now included the canonical roles of these four septin-associated kinases in the Introduction section of the revised manuscript.

      Please refer to line numbers 79-85.

      We have also revised sections describing corroborative findings and explicitly cited previous studies wherever relevant in the Results section to better distinguish previously established observations from the novel findings presented in this work.

      Other minor comments:

      (1) In Figure 1D-1F, also show the data for hls1∆ and kcc4∆ - which are shown in S3AC in the current version:

      We have now included the Inn1-mNG residence time in hsl1Δ and kcc4Δ cells, alongside elm1Δ and gin4Δ in Fig. 1E.

      Please refer to Fig. 1E in the revised manuscript.

      (2) The authors should be more careful in interpreting their negative Y2H data in Figure 2F.

      We have now explicitly discussed the caveats and inherent limitations of the Yeast Two-Hybrid assay in the manuscript.

      Please refer to line numbers 248-255.

      (3) Please provide quantification for Figure 5A, Figure S6I:

      We have added quantitative analyses showing rescue of Cdc3-mCherry and Myo13xmCherry mislocalization in gin4Δ and gin4Δ hsl1Δ strains upon artificial tethering of Elm1-GFP to the bud neck via Shs1-GBP.

      Please refer to Figures 8E, 8F, and S10B in the revised manuscript.

      (4) In Figure S6: label is missing "∆" in front of hsl1:

      We thank the reviewer for pointing out this error. We have corrected the labels accordingly.

      (5) As common consensus on yeast gene nomenclature, I suggest the use of "gene∆" instead of "∆gene":

      We thank the reviewer for this suggestion. We have now consistently placed the Δ symbol after deleted gene names throughout the manuscript in accordance with standard yeast nomenclature.

      (6) Lines (535-536): min(distribution) and max(distribution) in the formula is confusing. Clarify it or if possible use "minimum value", "maximum value" instead:

      We thank the reviewer for this suggestion. We have replaced the term “distribution” with “value” in the formula for protein accumulation kinetics analysis.

      Please refer to the updated formula in the Materials and Methods section (Lines 752753).

      (7) In line 86, "Dynamics of Septin-associated kinases and their defects during the cell cycle": Change the title as it is not clear what is meant by "their defects" given the discussed results under this title:

      We thank the reviewer for this suggestion. We have revised the section heading from “and their defects during the cell cycle” to “and defects associated with their deletions”.

      Please refer to line numbers 99-100.

      (8) On the Hof1-mNG image (Fig2C), show the line used for the line scan profile. Additionally, a similar line-scan profile could be useful in Figure S3G:

      As suggested by Reviewer 3, we removed the line-scan analysis from Figure 2 in the revised manuscript because Hof1 ring organization showed substantial heterogeneity across cells, making line-scan analysis difficult to interpret reliably.

      Reviewer #3 (Recommendations for the authors):

      Major points:

      (1) The % phenotype units are terrible. With no explanation, we do not really know whether they represent the percentage of cells that have a particular phenotype, or whether they correspond to a metric that measures some deviation between normal and extreme phenotypes. I would strongly recommend using precise quantitative metrics systematically (i.e. intensities, aspect ratios, division times, etc.) to properly quantify phenotypes:

      We thank the reviewer for suggesting the inclusion of precise quantitative metrics to assess phenotypic differences in the GFP-GBP tethering experiments. In the revised manuscript, we adopted quantification workflows that have been extensively validated and widely used in the literature, including those reported by Marquardt et al., 2024 (Fig. 7B and 7D) [2] from the Bi Lab. In response to the reviewer’s suggestion, we have now incorporated additional quantitative measurements, including cell area and aspect ratio (defined as the ratio of the cell’s major axis to the minor axis), for the experimental datasets presented in the manuscript.

      In the main figures, we now include aspect ratio quantification, while additional parameters are provided for the reviewer’s reference. We also quantified the fluorescence intensity of tethered proteins at the large bud neck and present these data together with the aspect ratio analysis in Figure 4 for elm1Δ cells in which Gin4GFP is artificially tethered to the bud neck via Hsl1-GBP. These quantitative analysis corroborates our qualitative observations and further strengthens our conclusions. Please refer to Figures 4D and 4E as representative examples.

      We have also changed the y-axis labels throughout the revised manuscript. For example, the y-axis in Fig. 7B is now labelled as “Cells exhibiting round morphology (%)”. Please refer to Fig. 4H, 4K, 7C, 8D, S5G, S6C, S7C, S9G and S9J for inclusion of aspect ratio quantification. Statistical analyses for the represented graphs were performed using Kruskal-Wallis nonparametric test, (N=3, n>150 cells/strain) (*: p<0.05, **: p<0.01, ****: p<0.0001, ns: p>0.05).

      Author response image 2.

      (2) Some quantitative analyses were performed manually where simple automated analysis should be performed to provide unbiased, accurate quantification:

      We fully agree with the reviewer that automated image analysis approaches, such as segmentation-based methods, are generally preferred for minimizing bias in morphological quantification. However, elm1Δ and gin4Δ cells exhibit severe phenotypes, including pronounced elongation and clumping, which makes reliable automated segmentation technically challenging for accurate quantification of parameters such as aspect ratio and cell area. For this reason, we used manual annotation for these analyses, as this approach enabled accurate delineation of individual cell and reliable measurements of morphological parameters such as cell area and size across the datasets despite being more time-consuming.

      Please refer lines 769-775 in the Materials and Methods section.

      (3) The "tethering" data also lack clear quantification. The authors should properly quantify the average intensity of Hsl1-GFP at the bud neck in each condition and correlate the results with cell aspect ratios or any other relevant parameters. For example, when comparing elm1null and elm1null Bud4-GBP with elm1null Kcc4-GBP, the visual impression is that as much Hsl1-GFP protein is recruited to the bud neck, whereas the phenotypes are dramatically different:

      We thank the reviewer for pointing this out. We have revised the image representation to facilitate clearer interpretation of the tethering experiments. In addition, we performed the key GFP-GBP tethering experiments using GBP-ymScarletI constructs, allowing direct visualisation of both the GFP-tagged protein and the GBP-tagged partner at the bud neck following tethering. We have included quantitative analyses of cellular morphology, raw fluorescence intensities of GFP-tagged proteins at the large bud neck, and the corresponding aspect ratio measurements for these updated datasets (Author response images 3, 4, 5). We have also included a summary table (Author response table 1) compiling these quantitative results for easier comparision.

      (4) I am very confused by Figure 3 which shows normal localization of Gin4-GFP in elm1null cells and seems to contradict other claims in the manuscript. This is very problematic for the interpretation of most of the "tethering" data:

      We understand the reviewer’s concern and have replaced the representative images of Gin4-GFP in elm1Δ cells in Figure 4B. Although Gin4-GFP is initially recruited to the presumptive bud neck during bud emergence in elm1Δ cells, it subsequently becomes mislocalized to the bud cortex during early cell cycle stages, resulting in reduced bud neck localization. As the cell cycle progresses, the Gin4-GFP signal at the bud neck decreases substantially in elm1Δ cells while remaining stable in wild-type cells until its departure prior to septin HDR remodelling (Fig. S5A-S5D).

      (5) Figure 6 is neither explained in the text nor in its legend. Could the authors explain the model and offer a comprehensive picture of the current knowledge?

      We have simplified the representative model to more clearly distinguish previously established knowledge from the findings presented in this study. Based on our results, We propose that Hsl1 functions both downstream of and in coordination with Elm1 and Gin4 to regulate septin stability and the timely execution of cytokinesis. Deletion of Elm1 disrupts the normal localization and crosstalk between Gin4 and Hsl1 at the bud neck, leading to septin mislocalization and misregulation of AMR dynamics, thereby revealing a previously uncharacterized role for Hsl1 in cytokinesis. The Results section has also been updated to reflect the revised model.

      Please refer to Figure 9 and lines 520-536.

      (6) Could the authors provide information about the double/triple mutant kinase phenotypes to clarify the overlap of functions among them?:

      Barral et al., 1999 [6] reported that individual deletions of Hsl1 and Gin4 result in mild cytokinetic defects, whereas deletion of Kcc4 does not produce any striking phenotype compared to wild-type cells. In contrast, the hsl1Δ gin4Δ kcc4Δ triple mutant remains viable but exhibits severe morphological abnormalities, including branched chains of elongated cells with defective cell separation. These mutants also display aberrant septin organization at the bud neck, characterized by irregular patch-like structures. Analysis of double mutants (hsl1Δ gin4Δ, gin4Δ kcc4Δ, and hsl1Δ kcc4Δ) revealed intermediate phenotypes between the corresponding single and triple mutants, with the hsl1Δ gin4Δ combination showing the strongest defects. Together, these findings suggest that the Nim1-related kinases function redundantly to regulate Swe1 activity and maintain septin architecture at the bud neck.

      Further supporting this model, Bouquin et al., 2000 [4] demonstrated that Elm1 operates independently of the Nim1-related kinases in controlling septin organization. The hsl1Δ gin4Δ kcc4Δ elm1Δ quadruple mutant exhibits severe septin localization defects and strong growth defects, in contrast to the elm1Δ single mutant, which primarily displays septin mislocalization from the bud neck to the bud cortex. These findings indicate that the combined activity of these kinases is essential for proper septin anchorage at the division plane and for assembly of the septin ring.

      Importantly, the progressively stronger phenotypes observed in double, triple and quadruple mutants also suggest that these kinases retain partially specialized functions at the bud neck. Consistent with this framework, our results support a model in which Nim1-related kinases function redundantly to regulate septin architecture and cytokinesis, likely through modulation of the AMR machinery. Because the localization of these kinases appears interdependent, as reported previously (Marquardt et al., 2020; Marquardt et al., 2024) [2,7] and corroborated by our findings, interpretation of mutant phenotypes remains complex and future studies will be required to delineate their individual contributions more precisely.

      Minor points:

      (1) Abbreviations are not defined in the manuscript.

      We have now expanded and defined all abbreviations throughout the manuscript.

      (2) Some of the writing in the figures is too small. Please make sure that a minimal size of letters/numbers is respected:

      We thank the reviewer for raising this issue. We have enlarged the figure labels and axis labels throughout the revised manuscript to improve readability.

      (3) Figure 2D. I am not sure that the line scans bring any useful information as the rings in mutant cells are quite heterogenous. This panel is also not cited in the text. Please make sure that every panel is cited at least once:

      We agree with the reviewer regarding the heterogeneity observed in the Hof1 ring organization in mutant cells and have therefore removed the line-scan analysis from the revised manuscript.

      (4) Knocked-out genes are written incorrectly. Please use the usual yeast nomenclature:

      We thank the reviewer for this suggestion. We have corrected the nomenclature for all deleted genes throughout the manuscript in accordance with standard yeast nomenclature.

      Author response image 3.

      Artificial tethering of Gin4-GFP to the bud neck via Hsl1-GBP-ymScarletI rescues cellular morphology in elm1Δ cells. (A) Representative images showing artificial tethering of Gin4-GFP via Shs1-GBP-ymScarletI, Hsl1-GBP-ymScarletI, Bud4-GBPymScarletI and Bni5-GBP-ymScarletI in elm1Δ cells. DC*=Differential contrast. Scale bar5µm. (B) Bar graph representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), one-way ANOVA Tukey’s multiple-comparison test (**: p<0.01, ****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=397, elm1Δ: n=408, elm1Δ-Shs1-GBPymScarletI: n=467, elm1Δ-Hsl1-GBP-ymScarletI: n=462, elm1Δ-Bud4-GBP-ymScarletI: n=401 and elm1Δ-Bni5-GBP-ymScarletI: n=317 cells). (C) Quantification of aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001, ns: p>0.05), (N=3, n>170 cells/strain). (D) Graph depicting the raw fluorescence intensity of Gin4-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=166, elm1Δ: n=177, elm1Δ-Shs1-GBP-YmScarletI: n=176, elm1Δ-Hsl1-GBPymScarletI: n=188, elm1Δ-Bud4-GBP-ymScarletI: n=185 and elm1Δ-Bni5-GBP-ymScarletI: n=151 cells).

      Author response image 4.

      Artificial tethering of Hsl1-GFP to the bud neck via septins or Nim1-related kinases rescues cellular morphology in elm1Δ cells. (A) Representative images showing the relocalization of Hsl1-GFP to the bud neck in elm1Δ cells via Shs1-GBP-ymScarletI and Gin4GBP-ymScarletI. Scale bar-5µm. (B) Bar graph representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), one-way ANOVA Tukey’s multiple comparison test (****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=508, elm1Δ: n=418, elm1ΔShs1-GBP-ymScarletI: n=535 and elm1Δ-Gin4-GBP-ymScarletI: n=482 cells). (C) Quantification of the aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (*: p<0.05, ****: p<0.0001), (N=3, n>165 cells/strain). (D) Quantification of raw fluorescence intensity of Hsl1-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001), (N=3, wildtype: n=156, elm1Δ: n=166, elm1Δ-Shs1-GBP-ymScarletI: n=169 and elm1Δ-Gin4-GBP-ymScarletI: n=165 cells.

      Author response image 5.

      Targeted localization of Kcc4-GFP to the bud neck via Hsl1-GBP-ymScarletI rescues cellular morphology in elm1Δ cells. (A) Representative images showing artificial tethering of Kcc4-GFP to the bud neck in elm1Δ cells via Shs1-GBP-ymScarletI, Hsl1-GBPymScarletI and Gin4-GBP-ymScarletI. Scale bar-5µm. (B) Quantitative analysis representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), oneway ANOVA Tukey’s multiple-comparison test (****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=443, elm1Δ: n=489, elm1Δ-Shs1-GBP-ymScarletI: n=312, elm1Δ-Hsl1-GBP-ymScarletI: n=563 and elm1Δ-Gin4-GBP-ymScarletI: n=337 cells). (C) Quantification of the aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (****: p<0.0001, ns: p>0.05), (N=3, n>165 cells/strain). (D) Quantification for the raw fluorescence intensity of Kcc4-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (**: p<0.01, ***: p<0.001, ****: p<0.0001), (N=3, wildtype: n=163, elm1Δ: n=168 elm1Δ-Shs1-GBP-ymScarletI: n=161, elm1Δ-Hsl1-GBP-ymScarletI: n=172 and elm1Δ-Gin4-GBP-ymScarletI: n=150 cells).

      Author response table 1.

      Summary table showing rescue of elongated morphology in elm1Δ cells upon forced recruitment of Nim1-related kinases via septins or its related kinases tagged with GBPymScarletI.

      Additional changes:

      The graph in Fig. S3D (revised preprint) has been updated to reflect a slight increase in the Chs2-mNG residence time in both the elm1Δ and gin4Δ strains, whereas our previous version indicated a delay only in the gin4Δ strain. Because the elm1Δ strain exhibited a more pronounced phenotype than the gin4Δ strain, we re-examined the analysis. The revised results show that the residence time of Chs2 during cytokinesis is modestly prolonged by approximately 2 minutes in both backgrounds. Accordingly, the graph and statistical analyses have been updated.

      References:

      (1) Mortensen, E.M., McDonald, H., Yates, J., and Kellogg, D.R. (2002). Cell Cycle-dependent Assembly of a Gin4-Septin Complex. Molecular Biology of the Cell 13, 2091-2105. 10.1091/mbc.01-10-0500.

      (2) Marquardt, J., Chen, X., and Bi, E. (2024). Reciprocal regulation by Elm1 and Gin4 controls septin hourglass assembly and remodeling. J Cell Biol 223. 10.1083/jcb.202308143.

      (3) Asano, S., Park, J.E., Yu, L.R., Zhou, M., Sakchaisri, K., Park, C.J., Kang, Y.H., Thorner, J., Veenstra, T.D., and Lee, K.S. (2006). Direct phosphorylation and activation of a Nim1-related kinase Gin4 by Elm1 in budding yeast. J Biol Chem 281, 2709027098. 10.1074/jbc.M601483200.

      (4) Bouquin, N., Barral, Y., Courbeyrette, R., Blondel, M., Snyder, M., and Mann, C. (2000). Regulation of cytokinesis by the Elm1 protein kinase in Saccharomyces cerevisiae. Journal of Cell Science 113, 1435-1445. 10.1242/jcs.113.8.1435.

      (5) Okada, H., MacTaggart, B., and Bi, E. (2021). Analysis of local protein accumulation kinetics by live-cell imaging in yeast systems. STAR Protoc 2, 100733. 10.1016/j.xpro.2021.100733.

      (6) Barral, Y., Parra, M., Bidlingmaier, S., and Snyder, M. (1999 Jan 15). Nim1-related kinases coordinate cell cycle progression with the organization of the peripheral cytoskeleton in yeast. Genes & Development 13. 10.1101/gad.13.2.176.

      (7) Marquardt, J., Yao, L.L., Okada, H., Svitkina, T., and Bi, E. (2020). The LKB1-like Kinase Elm1 Controls Septin Hourglass Assembly and Stability by Regulating Filament Pairing. Curr Biol 30, 2386-2394 e2384. 10.1016/j.cub.2020.04.035.

    1. eLife Assessment

      This useful study provides new insights into the liver-stage antigen LSA3, its export to erythrocytes, and its role in liver-stage development. While the functional importance of LSA3 is well demonstrated, the data underlying the conclusions regarding antibody specificity, liver-stage localization, and phenotype remain incomplete. A key strength of the study is the use of mosquito and humanized mouse models to access life-cycle stages that are rarely studied in most laboratories.

    2. Reviewer #1 (Public review):

      Summary:

      The extent P. falciparum liver stage parasites export proteins into the host cell is unclear. Most blood stage exported proteins tested in liver stages were not exported. An exception is LISP2 that is exported in P. berghei but not P. falciparum liver stages. While the machinery for export is present in liver stages, efforts to demonstrate export have so far been mostly unsuccessful. Parasite proteins exported during the liver stage could be presented by MHC and thereby become the target of immune control, incentive to study liver stage export and identify proteins exported during this stage. However, particularly for P. falciparum it is very difficult to study liver stages.

      This work studies LSA3 in P. falciparum blood and liver stages. The authors show that this protein is exported into the host cell in blood stages but in liver stages no or only very little export was detected. A disruption of LSA3 reduced liver stage load in a humanized mouse model, indicating this protein contributes to efficient development of the parasites in the liver.

      The paper also studied the localization of LSA3 in blood stages and used a known inhibitor to show that it is processed by plasmepsin 5, a protease important for protein trafficking. The work also showed that LSA3 is not needed for passage through the mosquito.

      Strengths:

      The main strength of this work is the use of the humanized mouse model to study liver stages of P. falciparum, which is technically challenging and requires specialized facilities. The biochemical analysis of LSA3 localization and processing by plasmepsin 5 are thorough and mostly overcame adverse issues such as a cross-reactive antibody and negative influence of the GFP-tag on LSA3 trafficking. The mosquito stage analysis is also notable as these kinds of studies are difficult with P. falciparum. However, there was no evidence for a function of LSA3 in mosquito stages.

      Weakness:

      The cross-reactivity of the antibody together with the co-infection strategy prevents reliable assessment of LSA3 localization in liver stages. Despite of this it seems LSA3 is not exported in liver stages and the paper does not bring us closer to the original goal of finding an exported liver stage protein.

      While the localization analysis in blood stages is well done and thorough, the advance is somewhat limited. LSA3 may be in structures like J dots, but this hypothesis was not tested. Although parasites with a disrupted LSA3 were generated, the function of this protein was not explored. However, this was now done in a separate study focussing on blood stage parasites (PMID: 41135800).

      Due to the difficulty of working with humanised mice, it was not possible to refine some of the conclusions and questions remain:<br /> The impact on liver stage development is interesting, but which phase of the liver stage is affected, and the phenotype remain largely unknown. The co-infection used (WT together with LSA3 mutant) has the advantage of a direct comparison of the mutant with the control in the same liver but complicates phenotypic analysis if the LSA3 antibody is also cross-reactive in liver stages. This issue adds a question mark to the shown localization and precludes phenotypic comparisons. It was also not possible to determine if the cross-reactive protein is expressed at that stage. While this might have been evident from the mixed WT/mutant infection (if all cells are positive for LSA3 there is cross-reaction; if about half of the cells are negative, there isn't) but assessing this failed.

      Significance:

      It is important information that LSA3 contributes to efficient liver stage development. However, neither LISP2 nor LSA3 seem to be exported in P. falciparum liver stages and can't confirm the potential of vaccines with proteins exported in this stage. LSA3 is still important and may still be the target of the immune response, but based on this work, probably not due to export in liver stages.

    3. Reviewer #2 (Public review):

      Summary:

      Immunogenic Plasmodium falciparum proteins that could be targeted to prevent parasite development in the liver are of significant interest for novel anti-malarial vaccine development. In this study, McConville et al evaluate the trafficking and functional importance of LSA3, a protein expressed in the blood and liver stages and previously shown to provide protection in immunized chimpanzees. LSA3 contains a PEXEL motif but the authors have previously shown that this protein does not appear to be exported beyond the PVM in the liver stage (McConville et al PNAS 2024). However, LSA3 trafficking and functional importance have not been comprehensively evaluated across stages. In the present study, the authors find that blood-stage LSA3 undergoes PEXEL processing and a portion of the protein is exported into the erythrocyte where it localizes to punctate structures distinct from Maurer's clefts. Using a knockout mutant, LSA3 is shown to be dispensable for blood and mosquito stages but important to liver-stage development. Collectively, these results validate LSA3 as a liver-stage target and place it among several other PEXEL proteins that display differential trafficking beyond the PVM in the erythrocyte but not the hepatocyte.

      Strengths:

      (1) The authors present a thorough analysis of LSA3 trafficking in the blood stage. PEXEL processing by Plasmepesin 5 is clearly demonstrated through a combination of mini LSA3-GFP reporters and Plasmepsin 5 inhibitors. Importantly, an LSA3 knockout mutant is used to show that the LSA3-C anti-sera also reacts with additional, unidentified parasite proteins in the blood stage. Nonetheless, comparison between the WT and KO parasites clearly indicates that a portion of LSA3 is exported into the erythrocyte, which is further supported by protease-protection assays with fractionated iRBCs. This contrasts with the liver stage where LSA3 does not appear to traffic beyond the PVM, similar to what has been observed for other PEXEL proteins in the rodent malaria model.

      (2) This study provides the first analysis of LSA3 exoerythrocytic function, showing this protein is important for liver stage development in chimeric human liver mice. Several PEXEL proteins in P. berghei have been shown to be exported into the host cell in the blood stage but do not appear to cross the PVM in the liver stage. These observations reinforce that even without detectible export into the hepatocyte, PEXEL proteins play critical roles during liver stage development.

      Weaknesses:

      The authors previously reported that anti-LSA3-C signal in the liver stage localizes within the parasite and at the parasite periphery but is not exported into the hepatocyte. In the present study, it is shown that anti-LSA3-C reacts with other parasite proteins beyond LSA3 in the blood stage and this may also occur in the liver stage. However, since liver-stage IFAs were only performed on samples co-infected with both WT and ∆LSA3 parasites, non-specific anti-LSA3-C reactivity at this stage could not be determined and the localization of LSA3 in the liver stage remains somewhat unclear.

      Comment on revisions:

      The authors thoughtfully addressed the reviewer comments.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript provides a comprehensive characterization of the Plasmodium falciparum protein LSA3, combining biochemical, genetic, and in vivo approaches. The authors convincingly demonstrate that LSA3 is expressed during liver stage infection and that disruption of the gene leads to a modest but reproducible reduction in liver stage parasite load in humanized mice.

      Strengths:

      Their biochemical and cell biological analysis of blood stages provides strong evidence that LSA3 is exported to the infected erythrocyte, and the detailed analysis of its PEXEL motif processing is well executed.

      Weaknesses:

      The study suggests LSA3 as one of only two known P. falciparum PEXEL proteins contributing to this stage, although there is no evidence for the export beyond the vacuolar membrane. Several key conclusions, particularly regarding antibody specificity, localization in liver stage parasites, and the interpretation of the phenotypic data, are not fully supported by the current experiments.

      Comments on revised version.

      I appreciate the authors' efforts to revise the manuscript and to clarify several aspects of the study.

      However, I remain concerned that some conclusions extend beyond the data presented. In particular, the authors acknowledge in their rebuttal letter that antibody specificity in liver stages could not be validated and that cross-reactivity cannot be excluded. Consequently, the localization data shown in Figure 5 cannot currently be considered definitive evidence for liver stage localization of LSA3 itself.

      Similarly, the revised manuscript appropriately states that LSA3 was not detected beyond the PVM in liver stages and that export into the hepatocyte remains unresolved. Nevertheless, several statements continue to imply a role for liver stage protein export. At present, the possibility that a domain of LSA3 may face the host-cell side of the PVM remains speculative and is not supported by direct experimental evidence.

      The liver stage fitness phenotype is convincing and supports the conclusion that LSA3 contributes to normal liver stage development. However, the current data do not establish the developmental process affected nor connect the phenotype to export beyond the PVM.

      I therefore recommend that the manuscript consistently distinguish between (i) demonstrated export of LSA3 during blood stage infection and (ii) the unresolved localization and trafficking of LSA3 during liver stage infection. I would also encourage the authors to consider revising the title to better reflect the findings presented, as the current title may be interpreted as demonstrating liver stage export, which has not been shown.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The extent to which P. falciparum liver stage parasites export proteins into the host cell is unclear. Most blood-stage exported proteins tested in liver stages were not exported. An exception is LISP2, which is exported in P. berghei but not P. falciparum liver stages. While the machinery for export is present in liver stages, efforts to demonstrate export have so far been mostly unsuccessful. Parasite proteins exported during the liver stage could be presented by MHC and thereby become the target of immune control, an incentive to study liver stage export and identify proteins exported during this stage. However, particularly for P. falciparum, it is very difficult to study liver stages.

      This work studies LSA3 in P. falciparum blood and liver stages. The authors show that this protein is exported into the host cell in blood stages, but in liver stages, no or only very little export was detected. A disruption of LSA3 reduced liver stage load in a humanized mouse model, indicating this protein contributes to efficient development of the parasites in the liver.

      The paper also studies the localization of LSA3 in blood stages and uses a known inhibitor to show that it is processed by plasmepsin 5, a protease important for protein trafficking. The work also shows that LSA3 is not needed for passage through the mosquito.

      Strengths:

      The main strength of this work is the use of the humanized mouse model to study liver stages of P. falciparum, which is technically challenging and requires specialized facilities. The biochemical analysis of LSA3 localization and processing by plasmepsin 5 is thorough and mostly overcame adverse issues such as a cross-reactive antibody and the negative influence of the GFP-tag on LSA3 trafficking. The mosquito stage analysis is also notable, as these kinds of studies are difficult with P. falciparum. However, there was no evidence for a function of LSA3 in mosquito stages.

      We thank the reviewer for their perspective on the strengths of the study.

      Weaknesses:

      The cross-reactivity of the antibody, together with the co-infection strategy, prevents reliable assessment of LSA3 localization in liver stages. Despite this, it seems LSA3 is not exported in liver stages, and the paper does not bring us closer to the original goal of finding an exported liver stage protein.

      While the localization analysis in blood stages is well done and thorough, the advance is somewhat limited. LSA3 may be in structures like J dots, but this hypothesis was not tested. Although parasites with a disrupted LSA3 were generated, the function of this protein was not explored. Given that a previous publication found some inhibitory effect of LSA3 antibodies on blood stage growth, a comparison of the growth of the LSA3 disruption clones with the parent would have been very welcome and easy to do. At this point, LSA3 is one more of many proteins exported in blood stages for which the function remains unclear.

      It might be possible to refine some of the conclusions. The impact on liver stage development is interesting, but which phase of the liver stage is affected, and the phenotype remains largely unknown. The co-infection (WT together with LSA3 mutant) has the advantage of a direct comparison of the mutant with the control in the same liver, but complicates phenotypic analysis if the LSA3 antibody is also cross-reactive in liver stages. This issue adds a question mark to the shown localization and precludes phenotypic comparisons. The authors write that they do not know if the cross-reactive protein is expressed at that stage. But this should be immediately evident from the mixed WT/mutant infection. If all cells are positive for LSA3, there is a cross-reaction. If about half of the cells are negative, there isn't. In the latter case, the localization shown in the paper is indeed LSA3, and morphological differences between WT and LSA3 disruption could be assessed without additional experiments.

      We thank the reviewer for their comments. While the LSA3-C antibody may cross-react with another parasite protein(s) in addition to binding LSA3 itself, we observed no strong evidence that this antibody localized beyond the liver-stage PVM, indicating that LSA3 is likely not targeted to the host cell compartment. We cannot exclude the possibility that a domain of LSA3 faces the hepatocyte lumen from this membrane and thus may be considered exported though follow-up studies are required (and are very challenging) to answer it. The phenotype of the NF54 DLSA3 mutant generated in this study at the blood stage was underway (by an independent lab in collaboration with us) and we are happy to disclose that the outcomes were recently published (May 2026) in an accompanying manuscript (PMID: 41135800). We completely agree that independently infected humanized mice would be helpful to address further remaining questions around the localization and temporal phenotype for LSA3 essentiality, which again will require follow up studies. In the present study, we intended to address whether LSA3 is important functionally, as this had not been reported.

      Significance:

      The conclusion from the paper that "our study presents just the second PEXEL protein so far identified as important for normal P. falciparum liver-stage development and confirms the hypothesized potential of exported proteins as malaria vaccine candidates" is partially misleading. Neither LISP2 nor LSA3 seems to be exported in P. falciparum liver stages, and we can't confirm the potential of vaccines with proteins exported in this stage. LSA3 is still important and may still be the target of the immune response, but based on this work, probably not due to export in liver stages.

      We thank the reviewer for the comment. We would like to emphasize the possibility that proteins localized at the PVM may be considered exported ‘if’ part or all of the protein (eg, a domain) faces the host cell lumen from the hepatocyte. We have not shown this to be the case for LSA3 or LISP2 but that possibility remains open. Nonetheless, LISP2 is exported (by P. berghei liver stages) and LSA3 is exported (by P. falciparum blood stages); both are exported proteins.

      Reviewer #2 (Public review):

      Summary:

      Immunogenic Plasmodium falciparum proteins that could be targeted to prevent parasite development in the liver are of significant interest for novel anti-malarial vaccine development. In this study, McConville et al evaluate the trafficking and functional importance of LSA3, a protein expressed in the blood and liver stages and previously shown to provide protection in immunized chimpanzees. LSA3 contains a PEXEL motif, but the authors have previously shown that this protein does not appear to be exported beyond the PVM in the liver stage (McConville et al, PNAS 2024). However, LSA3 trafficking and functional importance have not been comprehensively evaluated across stages. In the present study, the authors find that blood stage LSA3 undergoes PEXEL processing, and a portion of the protein is exported into the erythrocyte, where it localizes to punctate structures distinct from Maurer's clefts. Using a knockout mutant, LSA3 is shown to be dispensable for blood and mosquito stages but important to liver-stage development. Collectively, these results validate LSA3 as a liver-stage target and place it among several other PEXEL proteins that display differential trafficking beyond the PVM in the erythrocyte but not the hepatocyte.

      Strengths:

      The authors present a thorough analysis of LSA3 trafficking in the blood stage. PEXEL processing by Plasmepsin 5 is clearly demonstrated through a combination of mini LSA3-GFP reporters and Plasmepsin 5 inhibitors. Importantly, an LSA3 knockout mutant is used to show that the LSA3-C anti-sera also react with additional, unidentified parasite proteins in the blood stage. Nonetheless, comparison between the WT and KO parasites clearly indicates that a portion of LSA3 is exported into the erythrocyte, which is further supported by protease-protection assays with fractionated iRBCs. This contrasts with the liver stage, where LSA3 does not appear to traffic beyond the PVM, similar to what has been observed for other PEXEL proteins in the rodent malaria model.

      This study provides the first direct analysis of LSA3 function by reverse genetics, showing this protein is important for liver stage development in chimeric human liver mice. Several PEXEL proteins in P. berghei have been shown to be exported into the host cell in the blood stage, but do not appear to cross the PVM in the liver stage. These observations reinforce that even without detectable export into the hepatocyte, PEXEL proteins play critical roles during liver stage development.

      We thank the reviewer for their feedback regarding the strengths of the paper. 

      Weaknesses:

      A previous study reported that anti-LSA3 antibodies inhibit blood-stage growth, suggesting a role for LSA3 during erythrocyte infection. While the authors carefully evaluate the LSA3 mutant in mosquito and liver stages, the impact on blood stage fitness is not tested. While the knockout shows LSA3 is not essential in the blood stage, its importance during erythrocyte infection remains unclear.

      The authors previously reported that anti-LSA3-C signal in the liver stage localizes within the parasite and at the parasite periphery but is not exported into the hepatocyte. In the present study, it is shown that anti-LSA3-C reacts with other parasite proteins beyond LSA3 in the blood stage, and this may also occur in the liver stage. However, since liver-stage IFAs were only performed on samples co-infected with both WT and ∆LSA3 parasites, non-specific anti-LSA3C reactivity at this stage could not be determined, and the localization of LSA3 in the liver stage remains somewhat unclear.

      We thank the reviewer for their comments. The phenotype of the NF54 DLSA3 mutant generated in this study at the blood stage was underway (by an independent lab in collaboration with us) and we are happy to disclose that the outcomes were recently published (May 2026) in an accompanying manuscript (PMID: 41135800). While the LSA3-C antibody may cross-react with another parasite protein(s) in addition to binding LSA3 itself, we observed no strong evidence that this antibody localized beyond the liver-stage PVM, indicating that LSA3 is likely not targeted to the host cell compartment. We cannot exclude the possibility that a domain of LSA3 faces the hepatocyte lumen from this membrane and thus may be considered exported though follow-up studies are required (and are very challenging) to answer it. We completely agree that independently infected humanized mice would be helpful to address further remaining questions around the localization and temporal phenotype for LSA3 essentiality, which again will require follow up studies. In the present study, we intended to address whether LSA3 is important functionally, as this had not been reported.

      Reviewer #3 (Public review):

      Summary:

      This manuscript provides a comprehensive characterization of the Plasmodium falciparum protein LSA3, combining biochemical, genetic, and in vivo approaches. The authors convincingly demonstrate that LSA3 is expressed during liver stage infection and that disruption of the gene leads to a modest but reproducible reduction in liver stage parasite load in humanized mice.

      Strengths:

      Their biochemical and cell biological analysis of blood stages provides strong evidence that LSA3 is exported to the infected erythrocyte, and the detailed analysis of its PEXEL motif processing is well executed.

      We thank the reviewer for their comments.

      Weaknesses:

      The study suggests LSA3 as one of only two known P. falciparum PEXEL proteins contributing to this stage, although there is no evidence for the export beyond the vacuolar membrane. Several key conclusions, particularly regarding antibody specificity, localization in liver stage parasites, and the interpretation of the phenotypic data, are not fully supported by the current experiments.

      We understand the reviewer’s points. We agree that there is no evidence provided that LSA3 is targeted beyond the PVM; whether any of the protein faces the hepatocyte cytosol is unknown (and challenging to conduct) but this possibility remains plausible. LISP2- and LSA3deficient liver stages are less fit than parental controls and thus we stand by the conclusion that they are the two so far identified P. falciparum PEXEL proteins that are important for liver-stage development.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 163 says: "Altogether, this demonstrates that LSA3 is important but not critical for blood stage growth of P. falciparum": this is based on the cited Morita et al., 2017. However, previously LSA3 was considered dispensable based on a knock out in 3D7 (Maier et al., 2008; PMID: 18614010). Given that the authors generated a mutant for this work, it would be straightforward to test growth and clarify the importance of LSA3 in blood stages. If important, the analysis of the location and transport of LSA3 in blood stages would immediately become more relevant.  Maybe the data for this is already in the paper: the number of stage V gams was similar between mutant and control (Figure 4A). If this was calculated from the total number of asexual starting parasitemia, it includes blood stage growth, and it can be assumed that there is no growth defect in the mutant in the blood stages. If the number of stage 5 gams was calculated from the number of committed schizonts/rings, nothing can be said about blood stage growth, and asexual blood stage growth should be tested in specific experiments.

      We thank the reviewer for raising the function of LSA3 in blood stages and agree it was an obvious omission, though for good reason - a separate, collaborative study was underway. While this eLife preprint was in revision, our accompanying manuscript on the blood stage was published, showing the characterization of our NF54 DLSA3 mutant during blood-stage growth (PMID:41135800). The findings are now summarized and the citation included in the revised version of this preprint.

      Manuscript line 105: "although, notably, functional characterization of lsa3 deletion mutants has not yet been reported to confirm an important function": at least in blood stages, it was reported to be dispensable, see above. The corresponding study (Maier et al., 2008, PMID: 18614010) could be cited in that context. 

      The citation of PMID18614010 and 39913589 have now been added and we thank the reviewer.

      (2) Some questions central to the conclusions of this paper remain because it was unclear whether the serum did indeed detect LSA3 in the liver or not. It would be easy to check if all cells from the WT/Mutant mix experiment show LSA3 signal (this would mean it cross-reacts) or if only about half are positive (the mutants would be negative if there is no cross-reaction). This would be important to mention for Figure 5 because, at present, it is not known that what is labeled by the LSA3-C antibody in these images is (only) LSA3. 

      We thank the reviewer for this point and completely understand. We did check this via microscopy of liver sections co-infected with LSA3 mutant and control liver-stage parasites as we shared the reviewers line of enquiry. Unfortunately we could not detect parasites without LSA3 signal at the 5-day post-infection time point. This type of analysis does sound straightforward on paper but in reality is more challenging owing to several factors i) identifying sufficient individual parasites in an entire liver by microscopy can be challenging and variable from lobe to lobe and mouse to mouse, ii) the number of parasites required for a meaningful statistical analysis is increased due to coinfection of the liver (see Figure 4B as an illustration of this), iii) day 5 is a rather late liver-stage time point and so if there was a growth defect the defective parasites may be very small or sparse, iv) we cannot exclude that the LSA3 antibody may cross-react at the liver-stage, v) definitive conclusions are thus challenging and we feel require individual co-infections to be clear in the future. Nonetheless, the detailed qRT-PCR analyses identify a significant reduction in DLSA3 parasite liver load on day 5, indicating this protein is important for the human malaria parasite’s growth within human hepatocytes.

      (3) It is also unclear which parasites were imaged in Figure 5. The text of the results states that NF54 liver stages were used, but later: "As we employed a co-infection strategy to assess the essentiality of LSA3 versus NF54 in mice, we could not perform IFAs on individually infected mice in this study to validate the specificity of LSA3-C at the liver-stage". The legend says NF54 sporozoites on day 5 post-infection were used. I suspect it was a WT/mutant mix, in which case the above applies, and in the absence of cross-reactivity, half of the cells should be LSA3-C negative. If this is not the case, the localization in the liver becomes dubious.

      We apologize for the confusion and have corrected this. In Figure 5, we utilized liver sections from NF54-infected humanized mice that were stored at -80 C from a previously published study (McConville et al, PNAS 2024). Ideally, we would validate the specificity of LSA3 antibodies at the liver-stage using liver sections containing only DLSA3 parasites however the number of mice available was limited and the samples available to us also contained the Control line for qRT-PCR analyses (the co-infection strategy). As mentioned above, we couldn’t distinguish between these two strains by IFA at the time point analysed and this precluded us unequivocally validating the LSA3-C specificity in the liver-stage; however it cannot be excluded that the signal observed at the PVM is indeed LSA3. We are currently focusing research efforts on obtaining more humanised mice to answer this.

      Minor:

      (1) Introduction: Before the part on the PEXEL motifs, there are almost no references; please add references for all statements.

      We have added references.

      (2) Figure 1B is unclear regarding which part of the gene was deleted. The system used would permit a complete gene deletion, but the homology flanks seem to be within LSA3. If parts of the gene are left, the 75 kDa on the western blots might be a degradation product arising from both the truncated and the full-length protein. Please clarify in the sketch exactly where the homology flanks are, with respect to the start and stop of the gene. 

      The LSA3 gene was disrupted using the flanks as shown. The DHFR selection cassette comprises its own promoter and terminator such that insertion into the coding sequence completely disrupts expression of the protein thereafter, including the C-terminus within which the LSA3-C antibody binds. The new LSA3-T antibody described in our recently published accompanying manuscript that binds more N-terminally than LSA3-C also does not label the truncated protein. The original 5’ and 3’ flanks used for integration of the disrupted LSA3 allele by double cross-over recombination were then looped out into the original knockout plasmid and this was negatively selected against using exogenous 5-fluorocytidine (5-FC) via the suicide gene cassette CDUP (cytosine deaminase and uracil phosphoribosyl transferase that also contains a 5’ promoter and 3’UTR terminating element) in the construct. These features should provide clarification and have now been indicated in the figure and legend.

      (3) Line 161: Replace was with were.

      Corrected.

      (4) Figure 2, 224: Why do the authors think LSA3 must be in the luminal leaflet of the PVM as opposed to the outer leaflet of the plasma membrane?

      Several pieces of evidence combined led us to this conclusion in Figure 2B. i) if LSA3 was on the outer PVM leaflet, it would be substantially degraded in the EQT Pellet + PK fraction but a substantial population remained insensitive to PK, indicating much of the total protein pool was protected by the PVM (and possibly the parasite membrane; PM), ii) yet saponin, which leaves the PM intact, allowed PK to access and almost completely degrade LSA3 (see Saponin Pellet + PK), indicating that a substantial population of LSA3-C is located inside the boundary of the PVM, and this is membrane associated as saponin did not liberate it, rather, it remained in the Saponin Pellet before PK was added, iii) the TX-100 Super fraction confirmed LSA3 is membrane associated, as more is present in the TX-100 Super than the Saponin Super fractions, iv) if LSA3 was inside the PM, the Saponin Pellet fraction should be resistant to PK (as was the case for the cross-reactive band indicated with a red asterisk) but LSA3 (green asterisk) in the Saponin Pellet was PK sensitive. Altogether, our best conclusion from these data is that LSA3 is likely to be PVM associated with the LSA-C-binding domain facing internal to the PV, and a fraction is also exported beyond the PVM into the erythrocyte.

      (5) Line 245: GFP core "derived from digestion of the reporter in the food vacuole, which confirmed it was secreted from the parasite". I wonder if the amount of GFP "core" really can be used as evidence for secretion, and its amount can be compared between experiments. Did the author quantify this for the full-length protein to get a proportion per sample?

      Use of GFP core to measure defects in P. falciparum GFP reporter secretion has been described previously (for example PMID:23387285 and 35906227). The comparison the reviewer asked for is an interesting and important question: however the control would be to compare the ratio of GFP core to uncleaved in the control lanes as well, which is not possible to do since the full-length protein is digested by plasmepsin V in the native PEXEL versions of the experiments (mLSA3-GFP in the first blot, Vehicle in the second blot) leaving no full-length protein to compare to. It stands to reason that inhibition of N-terminal processing results in less protein removal from the membrane (ER or COPII vesicle or PM) resulting in less secretion out of the parasite for retrograde transport to the food vacuole with cytostomal vacuoles (analogous to plasmepsin II). In the food vacuole, the chimeras are in normal cases digested by proteases back to the GFP core that is resistant to cleavage and evident as GFP core on the immunoblots (PMID:10775264 and 14709539 and 19055692 and 20130643). 

      (6) Figure 3 has the word plasmid in two lanes. In Figure 3E, amend the labelling of the blots.

      We apologize for the formatting error in converting the figures to PDF during the original submission and thank the reviewer for the suggestion. This has now been corrected.

      (7) Lines 266/271/284: "live IFAs", live immunofluorescence assay. Does this mean an antibody was given to living   parasites?

      The correct term is live microscopy and this has been corrected.

      (8) Does Figure 6A fit with the data in Figure 6B? It seems 6B has a milder phenotype than 6A.

      We thank the reviewer for the question. Yes the data directly correspond to each other and are represented in two ways: Panel A shows the qRT-PCR raw data for liver load of each parasite strain per humanized mouse using a scientific scale on the y-axis. Panel B shows that magnitude of the DLSA3 defect as a percentage of the total liver load per mouse:

      % total parasite liver load  = ( strain 1 or strain 2 liver load ) x100

      sum of strain 1 + strain 2 liver loads

      The intent of showing both data is to convey the correct magnitude of the difference in two ways to assist the reader in understanding the true defect, both are accurate and both are statistically significant. In revision we detected mislabelling of humanized mouse 2 and 3 in the original graphs that has now been corrected and we sincerely thank the reviewer for helping us identify this error.

      (9) Line 482: Please add references for this debate. 

      These have been added.

      Reviewer #2 (Recommendations for the authors):

      Major Comments: 

      (1) In general, the authors have taken care not to overstate conclusions from their study. Nonetheless, while not technically inaccurate, the title might misleadingly suggest LSA3 is exported in the liver stage (this was my initial impression on reading it until I looked at the data). I suggest the authors revise the title to avoid confusion by clarifying that export was only observed in the blood stage.

      We sincerely appreciate the reviewer’s point. As this article was posted as a preprint that has now been cited several times, we have carefully weighed the comment and in the end decided to retain the current title for the above reason.

      (2) While the ability to generate the ∆LSA3 parasites clearly shows that the protein is not essential in the blood stage, the impact on parasite fitness is never tested but simply assumed (for instance, in lines 163-164: "...this demonstrates that LSA3 is important...for blood-stage growth..."). Do the ∆LSA3 parasites have a fitness defect in the blood stage consistent with the previous GIA data that would support this claim? Since the rabbit anti-LSA3-C antibodies produced by Morita et al did not have GIA activity against the blood stage, it is possible that the GIA observed with the human and mouse antibodies might have been due to reactivity with a different protein. If ∆LSA3 does cause a fitness defect, it would be interesting to know if the endogenous GFP-tagged line, which alters protein trafficking/membrane association, also produces this effect.

      We agree with the reviewer and would like to clarify that this omission was not intended to create confusion but was by design, due to a separate collaborative study that was underway to address such questions. While this eLife preprint was in revision, our accompanying manuscript on characterising NF54 DLSA3 at the blood stage was published (PMID:41135800). The findings are now summarized and the citation included in the revised version of this eLife preprint. In sum, LSA3 is not critical for erythrocyte invasion but its deletion perturbs the rate and efficiency of merozoite invasion, at the step(s) of resealing of the PVM/host cell, resulting in aberrant accole forms that protrude from the infected erythrocyte.

      (2) Figure 1D: While the images are compelling and I don't doubt the claim that LSA3 is exported in the blood stage (also supported by the fractionation/Pk experiments), the authors should provide quantification of the difference in exported signal between the WT and ∆LSA3 parasites in these IFAs to rigorously support this conclusion. Also, please include details about how many independent experiments are represented by the microscopy data throughout the manuscript (Figures 1, 2, 3, and 5).

      We understand the reviewer’s request and wish to indicate that the export signal was absent in all cells infected with DLSA3 that was imaged. The microscopy performed was from n=2-3 experiments except for Figure 5 which was from n=1 humanized mouse per time point in which multiple EEFs from the liver were imaged. This has been indicated in the figure legends. 

      (3) Careful inspection of the z-series images in Figure 5A shows that most of the LSA3-C signal seen outside the PVM (beyond the boundary delineated by EXP1) is closely associated with DAPI puncta, suggesting these are merozoites. Together with the prominent gap in the EXP1 signal, this suggests the schizont has already ruptured. Thus, anti-LSA3-C signal beyond the PV seems best explained as coming from merozoites or other material released by PV rupture, not from export across the PVM, and this should be added to the text in place of comments about localization to PV extensions or potential export (lines 358-359, 422-423).

      We do appreciate the reviewer’s careful eye and caution and are in complete agreement. We have added the comment as requested.

      Minor Comments:

      (1) The authors may want to denote the disordered repeat region in the LSA3 schematic in Figure 1A that is mentioned in the text.

      We have added the residue boundaries of the predicted domain from AlphaFold into both the schematic and the text and included a link to the LSA3 pages in PlasmoDB and

      AlphaFold in the Methods section.

      (2) The authors use rabbit anti-LSA3-C antibodies previously generated by Morita et al. These polyclonal antibodies were raised against a recombinant C-terminal region of LSA3 (residues 750-1433), but the schematic in Figure 1A indicates the antibodies recognize a smaller region between residues 1154-1433. Please adjust the figure accordingly, or if this is not the same antiLSA3-C antibody reported by Morita, please provide details about its production.

      The figure is corrected.

      (3) The authors use Alphafold to identify a region of LSA3 with similarity to the substrate binding domain of DnaK, but the data is not shown. Please include the Alphafold prediction in supplementary figures and provide information about how the predicted structural homology was determined.

      We have added a link to the AlphaFold page for PF3D7_0220000 in the methods.

      (4) The schematic in Figure 1B indicates that the DHFR cassette was inserted at an internal site within the lsa3 gene. If this is the case, it seems possible that an N-terminal portion of the protein is still expressed, but I was unable to find details about the boundaries of the homology flanks to determine the precise insertion site. Please clarify the knockout strategy and indicate the specific insertion site.

      The LSA3 gene was disrupted using the flanks as shown. The DHFR selection cassette comprises its own promoter and terminator such that insertion into the coding sequence completely disrupts expression of the protein thereafter, including the C-terminus within which the LSA3-C antibody binds. The new LSA3-T antibody described in our recently published accompanying manuscript that binds more N-terminally than LSA3-C also does not label the truncated protein. The original 5’ and 3’ flanks used for integration of the disrupted LSA3 allele by double cross-over recombination were then looped out into the original knockout plasmid and this was negatively selected against using exogenous 5-fluorocytidine (5-FC) via the suicide gene cassette CDUP (cytosine deaminase and uracil phosphoribosyl transferase that also contains a 5’ promoter and 3’UTR terminating element) in the construct. These features should provide clarification and have now been indicated in the figure and legend.

      (5) Line 162: I think this should read "antibodies that react with LSA3 were...".

      Corrected.

      (6) Figure 1D: The merge with the transmitted light channel is missing for the third panel in the ∆LSA3 IFAs. Also, please define the scale bar length in the legend.

      Corrected.

      (7) Lines 744-746: The IFA fixation panel order description (top, bottom) in the Figure 2A legend is reversed from what is shown in the actual figure. Also, please define the scale bar length. 

      Corrected.

      (8) Lines 184-186: Since the fractionation/PK protection assays suggest most of LSA3 is in the PV, it would be interesting to know if the strong peripheral/PV signal observed in the PFA-fixed IFAs in Figure 2A is also present in the ∆LSA3 parasites, or is this non-specific? 

      Thank you for the suggestion. We agree this would be an interesting result to know but do not have the capacity at the present time.

      (9) Lines 219-225: It is unclear to me why these results are interpreted to suggest that the majority of LSA3 is peripherally associated with the luminal leaflet of the PVM. Wouldn't an integral membrane configuration in the PVM (with the C-terminus facing the host cytosol) or PPM (with the C-terminus facing the parasite cytosol) also account for the data? Adding a carbonate extraction would help clarify this point.

      Several pieces of evidence combined led us to this conclusion in Figure 2B. i) if LSA3 was on the outer PVM leaflet, it would be substantially degraded in the EQT Pellet + PK fraction but a substantial population remained insensitive to PK, indicating much of the total protein pool was protected by the PVM (and possibly the parasite membrane; PM), ii) yet saponin, which leaves the PM intact, allowed PK to access and almost completely degrade LSA3 (see Saponin Pellet + PK), indicating that a substantial population of LSA3-C is located inside the boundary of the PVM, and this is membrane associated as saponin did not liberate it, rather, it remained in the Saponin Pellet before PK was added, iii) the TX-100 Super fraction confirmed LSA3 is membrane associated, as more is present in the TX-100 Super than the Saponin Super fractions, iv) if LSA3 was inside the PM, the Saponin Pellet fraction should be resistant to PK (as was the case for the cross-reactive band indicated with a red asterisk) but LSA3 (green asterisk) in the Saponin Pellet was PK sensitive. Altogether, our best conclusion from these data is that LSA3 is likely to be PVM-associated with the LSA-C-binding domain facing internal to the PV, and a fraction is also exported beyond the PVM into the erythrocyte. If the question is whether LSA3 is an integral PVM protein, we agree that use of carbonate in the future would answer that question.

      (10) Figures 3D and E: There are some problems with some of the text wrapping in these panels.

      We apologise, this was a formatting issue as the manuscript was converted to PDF.

      We have corrected this error.

      (11) Line 422-423: In fact, the Z-sections shown in Figure 5 appear to indicate that the LSA3-C signal is predominantly located within the parasite, not at the PVM.

      We do appreciate the reviewer’s careful eye and caution and are in complete agreement. We have corrected the final conclusion to be more accommodating of this.

      (12) Lines 468-470: Since cross reactivity of anti-LSA3-C is substantial in the blood stage but was not defined in the liver stage by analysis of unmixed infections, how do the authors know that they were not observing ∆LSA3 parasites in their IFAs? I think what they mean here is that parasites lacking anti-LSA3-C reactivity were not observed, which is an important distinction.

      The reviewer is correct and this has been corrected.

      (13) Lines 478-479: The authors should also mention that the P. berghei PEXEL proteins evaluated in Fougere et al are exported in the blood stage, similar to LSA3. Moreover, other studies have shown something similar for additional endogenous PEXEL proteins or reporters in P. berghei (PMIDs 22329949, 26347246, 34956312).

      We have added the additional text regarding export into the infected erythrocyte and the reference to IBIS1.

      (14) Line 491: The data here don't support that LSA3 is "required" for liver stage development, only that it is important to it. Since the authors have not defined the cross-reactivity of anti-LSA3C in unmixed infections, it is not clear that ∆LSA3 parasites are arrested early in the liver stage, only that they show a reduced number of genome copies relative to the parental control. 

      We have amended the sentence to “required for normal liver stage development”.

      (15) Line 530: I think NGF54 should be NF54.

      Corrected.

      Reviewer #3 (Recommendations for the authors):

      (1) Antibody specificity in liver stage IFA experiments:

      The specificity of the anti-LSA3 antiserum (LSA3-C) used in liver stage IFA is not fully convincing. While the KO parasites were used effectively to validate specificity in blood stages, the same is not true for liver stages. 

      (a) It is essential to repeat IFA with ΔLSA3 parasites in liver stage infections to determine whether the observed PVM staining is truly specific.

      We appreciate the reviewer’s point, however at a cost of over $5000 per humanized mouse, we do not have the capacity to conduct this experiment at the present time. We highlight that, as the blood stage IFAs confirmed the specificity of LSA3-C for LSA3, the possibility remains open that LSA3 is specifically recognized at the PVM.

      (b) If the antibody is the same polyclonal serum used in Morita et al. (2017), why did the authors not employ a monoclonal antibody, which they presumably have access to and which would provide greater specificity? 

      We have included new data confirming that LSA3 is exported using LSA3-T, in addition to LSA3-C.

      (c) Given that rabbit antisera often show non-specific staining at the PVM in liver stage parasites, co-localization with PVM markers is not sufficient. Inclusion of the ΔLSA3 parasites in liver stage IFA is critical. It will also show whether there is any cross-reaction of the antiserum in liver stage parasites, as seen by IFA for blood stage parasites. 

      We thank the reviewer for their feedback.

      (d) To validate the serum further, the authors should infect HC-04 cells in vitro with GFP-LSA3 parasites and stain with LSA3-C to confirm overlap between the tagged protein and the antibody signal.

      We thank the reviewer for their feedback.

      (e) For higher-resolution co-localization, expansion microscopy - now commonly used even in malaria research - would substantially improve the analysis. 

      We thank the reviewer for their feedback.

      (2) The localization of LSA3 in this study differs notably from Morita et al. 2017, who reported localization to dense granules in merozoites and staining in ring-stage parasites at the PVM. 

      (a) The authors confirm DG localization, but they do not examine ring-stage parasites. They should include the IFA of ring stages to clarify whether they can replicate the previous findings.

      We thank the reviewer for their feedback.

      (b) Additionally, the differences in Western blot banding patterns between the two studies should be addressed. Do the authors have an explanation for these discrepancies? 

      We thank the reviewer for their feedback.

      (3) The authors report a ~40% reduction in liver parasite load using qPCR, which is statistically significant. However, this phenotype is modest and should not be interpreted as showing that LSA3 is essential.

      (a) Please avoid terms like "required" or "essential" and instead describe the protein as "contributing to normal development" or "influencing fitness."

      We have used the term “required for normal liver stage development”.

      (b) Since the authors generated liver sections, they should take advantage of these to quantify the number and size of liver stage parasites, which would help determine whether the phenotype reflects fewer infected cells or reduced parasite growth.

      We did check this via microscopy of liver sections, but all mice were co-infected with LSA3 mutant and control liver-stage parasites, as we shared the reviewers line of enquiry. Unfortunately we could not detect parasites without LSA3 signal at the 5 day post infection time point. This type of analysis does sound straightforward on paper but in reality is more challenging owing to several factors i) identifying sufficient individual parasites in an entire liver by microscopy can be challenging and variable from lobe to lobe and mouse to mouse, ii) the number of parasites required for a meaningful statistical analysis is increased due to coinfection of the liver (see Figure 4B as an illustration of this), iii) day 5 is a rather late liver-stage time point and so if there was a growth defect the defective parasites may be very small or sparse, iv) we cannot exclude that the LSA3 antibody may cross-react at the liver-stage, v) definitive conclusions are thus challenging and we feel require individual co-infections to be clear in the future. Nonetheless, the detailed qRT-PCR analyses identify a significant reduction in DLSA3 parasite liver load on day 5, indicating this protein is important for the human malaria parasite’s growth within human hepatocytes.

      (c) It would also be valuable to include IFA from singly infected ΔLSA3 livers (rather than co-infected), and possibly at earlier timepoints, to identify the developmental window affected.

      We agree it would be valuable.

      (4) The manuscript suggests that LSA3 may be exported beyond the PVM into the hepatocyte, based on a small number of peripheral puncta.

      (a) This claim is not convincingly supported by the data. The punctate signals shown in Figure 5 are weak and may rather reflect PVM extensions or TVN. In fact, one punctum even overlaps with the DAPI signal (figure 5, middle panel), which raises further doubt about the localization.

      We appreciate the reviewer’s careful eye and caution and have added the comment regarding DAPI.

      (b) Given the lack of KO controls in these liver stage IFAs, the authors should not describe LSA3 as "exported beyond the PVM". The language should be revised to reflect that the protein localizes predominantly to the PVM, and any extra-PVM signal remains unconfirmed and could be non-specific. 

      (c) This is especially important given the well-known tendency of rabbit antisera to produce background PVM staining in liver stage parasites. 

      Corrected.

      (e) In an earlier report (McConville et al, 2024, PNAS), they clearly state that LSA3 is NOT exported beyond the PVM. Actually, the staining in the previous report looks quite different from the images provided for Figure 5. The authors might wish to comment on this. 

      We thank the reviewer for their feedback.

      Minor comments:

      In some sections, the manuscript uses "exported" to refer to trafficking to the PVM. This terminology should be used more carefully and consistently, since "export" often implies translocation into the host cytosol

      We understand that export involves a protein localizing within the host cell and so protrusion through the PVM may also be considered exported, however, we have not confirmed this for LSA3 in liver stages.

    1. eLife Assessment

      This study offers an important set of literature-based parameter ranges for calibrating cardiac electromechanical models. The key scientific claims are largely supported by convincing evidence and carefully-documented methodologies, but a few specific issues remain only partially supported, in particular relating to the locking-free formulation, reporting of adequate details regarding the electrophysiological model (Eikonal approach), and the fact that physiological accuracy is predominantly achieved in a single chamber (the LV).

    2. Reviewer #1 (Public review):

      Summary:

      The study by Wang et al. investigates cardiac electromechanical modeling and simulation techniques, focusing on the calibration and validation of ventricular models according to ASME V&V40 standards. The researchers aim to calibrate model parameters to align with key biomarkers such as QRS duration and left ventricular ejection fraction and validate the model against independent measurements such as displacement and strain metrics. The authors also examine the impact of parameter variations on deformation, ejection fraction, strains and other biomarkers. The overarching aim of the study is to give credibility to the underlying computational electromechanics framework as a step towards the cardiac Digital Twin vision.

      Strengths:

      (1) The study presents a solid validation strategy for cardiac models based on independent data.

      (2) It integrates electrophysiological, mechanical, and hemodynamic biomarkers for sensitivity analysis and calibration.

      Weaknesses and Limitations:

      (1) Model Assumptions: The study relies on several simplified modeling assumptions that do not reflect the current state-of-the-art:

      a) Isotropic scaling of the ventricular mesh to generate an unloaded reference geometry.

      b) Simplified afterload and preload models that do not consistently capture the full range of physiological responses.

      c) Simplified epicardial boundary conditions.

      These limitations are appropriately acknowledged and discussed by the authors in a dedicated Limitations section.

      (2) Numerical Framework:

      a) The numerical framework used for the mechanical part of the model may be susceptible to locking effects that could contribute to artificially stiff and less contractile behavior. This is indicated by a ten-fold scaling of the peak active contractile force parameter relative to literature values and notable sensitivity of the model to the tissue compressibility parameter. While - as acknowledged by the authors - part of this can be attributed to simplified modeling choices, other comparable studies have not reported similar issues.

      b) The human electrophysiology model is not described in enough detail. Currently, it is not mentioned in the manuscript that an Eikonal model was used to compute activation times on the endocardial surface to be robust against coarse mesh resolutions.

      (3) Geometrical model and digital twin: The model presented combines anatomical data, electrical measurements, and physiological reference values from different individuals or population averages, rather than being derived from a single patient. The authors have appropriately moderated their claims in the revision, framing the work as a step towards the cardiac Digital Twin vision rather than asserting that the model itself constitutes a digital twin.

      (4) Calibration procedure: The description of the calibration procedure has been substantially improved in the revision. The authors now provide explicit rationale for each calibration step and clarify that the procedure targets multiple physiological biomarkers in sequence. Verification that the calibrated model produces physiological cellular dynamics, including intracellular calcium transients, is now provided. The revised manuscript also shows the simulated electrocardiogram alongside population reference ranges, which partially addresses the question of calibration quality. However, a direct comparison of the simulated electrocardiogram with the individual measured signal used for calibration is not provided, which would give a more stringent and direct assessment of how well that specific calibration target was achieved.

      Comments on revised version.

      The revision represents a genuine improvement. The calibration procedure is now more transparently described, physiological cellular dynamics are verified, the digital twin framing has been appropriately moderated to reflect a step towards that vision rather than a claim of having achieved it, and an expanded limitations section identifies where the framework falls short.

      Several of the concerns raised in the first round have been addressed, but some issues remain:<br /> The model still requires a ten-fold scaling of the peak active contractile force relative to literature values, and the imbalance between left and right ventricular output persists, along with non-physiological right ventricular pressures and ejection fraction. These are partly fundamental limitations of the current modelling approach that may not be fully resolvable within the scope of this paper, and the authors are to be credited for acknowledging them. However, they do constrain the conclusions that can be drawn about the credibility of the framework for reproducing healthy cardiac physiology.

      The population-averaged reference dataset and the calibration and validation framework remain contributions of value to the community, and the revised limitations section adds useful transparency about the current state of the art.

    3. Reviewer #2 (Public review):

      The authors present an interesting study on calibrating and validating a biventricular cardiac electromechanical model. This is an important contribution, but some questions remain about the quantitative validation and verification aspects of the study.

      Major comments:

      (1) The title and paper stress the importance of validation on several occasions. However, the actual validation performed is limited to the section in lines 427-439. Furthermore, it is entirely qualitative, making assessing the model's quality difficult. Most of the paper is focused on sensitivity analysis, which is also interesting but unrelated to validation. Can you include a quantitative comparison with deformation biomarkers? E.g., spatially quantify strain differences between simulation and in vivo data, or overlay the current configuration of the geometry with MRI in various views, and calculate a displacement error norm.

      (2) You mention the ASME V&V40 standards throughout your paper. Yet, you only address the "second V" validation, ignoring the "first V" verification. How did you ensure that your computational models are implemented correctly?

      (3) All parameters discussed in this publication are physical parameters. What is the sensitivity of your model outputs concerning computational parameters?

      Comments on revised version.

      The authors have addressed my prior comments

    4. Author response:

      The following is the authors’ response to the original reviews.

      We greatly appreciate the reviewers for their efforts in reviewing our manuscript. We highlight that the key contributions of our paper are to provide a framework for calibration and validation of high-fidelity cardiac electromechanical models based on a diverse compilation of clinical datasets, and that we provide one example of such an evaluation of our own baseline electromechanical model. The comments raised by the reviewers were chiefly focused on the second goal, which is our specific model and the outcomes of the evaluation process, rather than on the evaluation framework itself. As such, we have made improvements to our model implementation and to provide additional confidence in our specific modelling framework through this review process. Specifically, we have strengthened the verification component of this evaluation, provided additional quantitative measures, and included a more in-depth discussion of the remaining limitations in our modelling framework. We hope that the updated version of the manuscript and our efforts to improve it are well-received by our reviewers and editors, as well as by the modelling and simulation community at large.

      eLife Assessment

      This is a potentially important study that explores the relevant range of parameter values for calibration and validation of cardiac electromechanics in ventricular models. Although much of the work presented is solid, the evidence provided to support the authors' key scientific claims is incomplete, especially as it relates to the emphasis on standardized validation and verification approaches. Notably, the level of model personalization presented in this work falls short of the threshold for what could reasonably be called a "digital twin", even by the relatively relaxed standards that have emerged in computational physiology and related fields in recent years.

      We appreciate the eLife assessment for identifying the potential importance of our study. Regarding the threshold for 'digital twin', we note that a cardiac digital twin is envisioned as a patient-specific computational model of the heart, personalised from multi-modal clinical data and continuously updated to support diagnosis, prognosis, and treatment planning, which is a goal that to our knowledge no published electromechanical study has simultaneously fulfilled. It is for this reason that the community refers to the 'digital twin vision' rather than its realisation, and we adopt this framing consistently throughout the manuscript.

      The primary contribution of this manuscript is the framework: a systematic application of ASME V&V40 standards to a fully coupled electromechanical model, spanning electrical, mechanical, and haemodynamic biomarkers within a single study. In our revision, we have clarified that the model evaluation presented here is an example application of that framework, which was designed not to certify a model as complete, but to provide a transparent audit of current capability that identifies where confidence is established and where further development is needed. We have updated the title and language throughout the manuscript to reflect this framing consistently.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study by Wang et al. investigates cardiac electromechanical modeling and simulation techniques, focusing on the calibration and validation of ventricular models according to ASME V&V40 standards. The researchers aim to calibrate model parameters to align with key biomarkers such as QRS duration and left ventricular ejection fraction, and validate the model against independent measurements such as displacement and strain metrics. The authors also examine the impact of parameter variations on deformation, ejection fraction, strains, and other biomarkers. The overarching aim of the study is to give "credibility to the underlying computational electromechanics framework" and to "pave the way towards credible cardiac electromechanical Digital Twins."

      Strengths:

      (1) The study presents a solid validation strategy for cardiac models based on independent data.

      (2) It integrates electrophysiological, mechanical, and hemodynamic biomarkers for sensitivity analysis and calibration.

      Weaknesses and Limitations:

      (1) Model Assumptions: The study employs simplified modeling assumptions that are not state-of-the-art, e.g.,

      (a) Isotropic scaling of the mesh to generate an unloaded reference geometry.

      (b) Simple afterload and preload models that fail to produce physiological results.

      (c) Simplified epicardial boundary conditions.

      While our model was able to broadly achieve physiological behaviour based on the calibration and validation datasets, it also contains several simplifications that can be expanded with more sophisticated techniques to allow explorations in specific areas. We have added a dedicated Limitations subsection to the Discussion section of the manuscript to address these and to provide references to relevant studies.

      (2) Numerical Framework:

      (a) The mesh resolution and/or the numerical framework used for the mechanical part appears to suffer from known numerical artifacts (locking effects), leading to overly stiff or inaccurate behavior in finite element analysis. This results in an artificially stiff response to deformation, which is compensated by setting active contraction to ten times the value reported in the literature. The authors attribute this to limitations in using ex vivo tissue measurements to represent in vivo function, although similar issues were not observed in previous works.

      We thank the reviewer for raising this point and have investigated it carefully. We have added a verification section as well as discussions to the manuscript to more comprehensively discuss this point. In short, through various tests against benchmark (Land) and comparing stress-strain curves in cube simulations with the same mesh resolution, we could not identify evidence of volumetric locking effects. The elevation in contractile force was also necessary in a simplified ellipsoid version of the model in a previous publication [ref 7, Levrero-Florencio, et al. 2020]. We note that these benchmarks were performed in the incompressible transversely isotropic regime; whether analogous locking effects exist in the dynamic orthotropic active contraction framework used in the full biventricular simulations remains an open question, which we have identified as a priority for future benchmarking, for example against the Arostica et al. 2025 benchmark.

      We agree that the explanation of this as ex vivo vs in vivo difference in contractile force is too simple, and other contributing factors are better understood through comparison with similar studies in the field. Strocchi et al. (2023) used a four-chamber model with explicit atrial mechanics, and in her history matching varied Tref within +- 33-55% of a reference value of 120-150 kPa, targeting a peak active tension of 160 +- 15 kPa, which was a considerably more modest adjustment than applied here, likely reflecting differences in model geometry, pericardial constraint, and circulatory model between the two studies. Gerach et al. (2021, Mathematics) applied manual parameter adjustments informed by in vivo active tension measurements of 120 – 150 kPa and achieved ejection fractions of approximately 63%; however, they reported that systolic pressures in both ventricles were too high for a healthy heart, and similarly reported elevated peak ejection rates compared to MRI measurements, a difficulty we also encountered. Notably, Gerach et al., report that atrial contraction contributes approximately 11-13% of end-diastolic volume, which in a biventricular-only model would directly reduce the achievable LVEF and necessitate compensating adjustments to active tension. Zingaro et al. (2024, Journal of Computational Physics), using an alternative active tension model (RDQ20), similarly found it necessary to increase contractility parameter (a_XB) to achieve sufficient ejection, and explicitly report that no single parameter configuration simultaneously achieved physiological peak ejection rate and LVEF, a fundamental tension we also encountered. Together, these comparisons suggest that the elevated Tref in our model most likely reflects a combination of the absence of atrial filling, simplifications in pericardial constraint, and the lack of poroelastic behaviour, rather than volumetric locking alone. Additional investigations are needed as explained in the manuscript, and these factors are identified as open priorities for future development within our framework.

      We added a comparison with these three studies to the discussion section of the manuscript and have updated the limitations section to reflect this more nuanced account of the factors contributing to the elevated active tension scaling. Further work will be required to address these points.

      (b) Further, the authors employ the monodomain model for the simulation of the electrical excitation and relaxation on a relatively coarse grid with an approximate edge length of 1mm. This resolution is known to be insufficient for reliable results in organ-scale electrophysiology modeling.

      Our ECG simulations are robust against coarse mesh resolutions since we use an Eikonal solution to prescribe the activation times on the endocardial surface and we tune the diffusivity parameters in the model such that the correct conduction velocities are reached, as performed in Camps et al. (2024) (ref 33) using the tuneCV tool in monoAlg3D (https://github.com/rsachetto/MonoAlg3D_C/tree/master/scripts/tuneCV), which is similar to the tool in openCARP (described here: https://opencarp.org/documentation/examples/02_ep_tissue/03a_study_prep_tunecv).

      While this mesh resolution may not be sufficient for simulations of more complex behaviour, such as re-entry and fibrillation patterns, it is sufficient for simulations of ECGs in this study. We have noted this in the methods section.

      (3) Geometrical model and digital twin: The geometrical model, taken from a public cohort and calibrated to an ECG of another individual along with population-averaged values from a databank (UK Biobank), and unrelated measurements from surgical procedures, can hardly be considered a digital twin. Further, validation of the model was then performed against data from yet another cohort.

      We thank the reviewer for this point and welcome the opportunity to clarify our dataset choices. The use of multiple data sources was a deliberate methodological decision. While an ideal dataset for electromechanical model evaluation would combine full biventricular geometry, 12-lead ECG, invasive pressure measurements, and myocardial strain data from a single individual, no such dataset currently exists in the public domain, and acquiring it routinely would be impractical in clinical settings. Multi-source integration therefore reflects the realistic deployment scenario for future clinical translation of these tools.

      The specific choice of geometry was principled: the mesh associated with the ECG dataset that was available to us was truncated at the base due to the clinical acquisition protocol, which would have prevented physiologically realistic basal boundary conditions. The female Rodero geometry we chose provides full ventricular coverage and was selected on that basis.

      Demonstrating that a coherent, systematically evaluated framework can be constructed from compiled multi-modal data is itself a contribution because it makes the tools accessible to the wider community without requiring a single ideally acquired dataset.

      (4) Calibration procedure: There are apparent flaws in the calibration procedure, or it is not described in sufficient detail. The authors dedicate significant effort to motivating parameter ranges, but in the end they use mostly other parameters for the calibration process, aiming to maximize left ventricular ejection fraction. It is not clear whether the chosen parameters result in, e.g., physiological calcium traces or calibrated parameters that are within physiological ranges.

      Thank you for raising this point, which we have now clarified in the manuscript. The parameters that were chosen for the calibration process were based on the results of the sensitivity analyses.

      In addition, we have supplemented results Figure 1 with a subfigure F showing that the calcium transient and action potential durations fall within physiological ranges after calibration.

      (5) Goodness of fits, e.g., a direct comparison of the measured and the simulated ECG, are not provided to assess calibration quality.

      The calibrated model achieves QRS duration of 89 ms and QT interval of 360 ms, both of which fall within the healthy reference ranges compiled in Table 2, providing a biomarker-level assessment of calibration quality (Figure 1A). A full quantitative goodness of fit analysis of the simulated ECG morphology was performed following the methodology of Camps et al. [52], in which the same beat-averaged ECG was processed; we direct the reader to that work for full details rather than reproducing the analysis here.

      (6) Due to these limitations and weaknesses, the authors fall short of achieving some of their goals, particularly establishing credibility for the underlying computational framework and in reproducing healthy pressure-volume loops, and in achieving physiological simulations while using physiological or reported ranges for the calibrated parameters.

      For example, a key physiological requirement is that the right and left ventricular stroke volumes are approximately equal in a heart beating at a limit cycle, as the blood pumped by the right ventricle into the pulmonary circulation must match the amount pumped by the left ventricle into the systemic circulation. This balance is not achieved in this study.

      We thank the reviewer for identifying the stroke volume imbalance. We acknowledge the physiological requirement that, in a steady-state limit cycle, the right ventricular stroke volume must approximately equal the left ventricular stroke volume. However, since our model does not explicitly prescribe volumes, to achieve this, we would need to either explicitly tune active tension for the left and right ventricles separately, such as done in https://www.frontiersin.org/journals/physiology/articles/10.3389/fphys.2021.716597/full or develop a more sophisticated circulatory model and employ a multistep procedure that sequentially tunes circulatory dynamics, passive mechanics, and active contraction, such as done in https://www.biorxiv.org/content/10.64898/2025.12.11.693778v1.full. Both of which are beyond the scope of this paper.

      We note that, despite the absence of explicit RV calibration, the RV volumetric measures and pressures remain within physiological ranges, suggesting that the coupled biventricular mechanics are broadly plausible. As such, we have noted this limitation in our discussion section, and sign-posted to other studies where the stroke volume match is achieved.

      (7) The conclusive claim that "the study paves the way towards credible electromechanical cardiac Digital Twins" is not supported. The model exhibits non-physiological behavior, requires unsupported parameter alterations (such as a 10-fold active stress scaling), and does not represent a digital twin, as model data are drawn from various unrelated, non-patient-specific sources.

      We thank the reviewer for this comment, which gives us the opportunity to clarify our use of the term 'digital twin'. A cardiac digital twin is envisioned as a patient-specific computational model of the heart, personalised from multi-modal clinical data and continuously updated to support diagnosis, prognosis, and treatment planning. This is a transformative goal for precision cardiology that the field is actively working towards, with credible, systematically validated electromechanical models as its essential foundation. To our knowledge, no published study in cardiac electromechanical modelling has simultaneously fulfilled all three requirements, and it is for this reason that the community often refers to the 'digital twin vision' rather than its realisation.

      The primary contribution of this manuscript is the framework: a systematic application of ASME V&V40 standards to a fully coupled electromechanical model, spanning electrical, mechanical, and haemodynamic biomarkers in a single study. The model evaluation presented here is an example application of that framework. Importantly, the framework is not designed to certify a model as complete, but to provide a transparent audit of current capability by identifying where confidence is established and where further development is needed. In this sense, the limitations surfaced through this evaluation are themselves a contribution: they define open problems and priorities for the field.

      We have added a definition of the digital twin concept and the roadmap towards its realisation to the introduction and have updated the language throughout the manuscript to consistently reflect the distinction between the framework contribution and the model evaluation. We maintain that this transparent approach represents a meaningful step towards the digital twin vision.

      The specific limitations of the current model implementation are addressed in detail in the relevant sections of this response and in the updated manuscript, where we have substantially strengthened the verification and discussion components.

      Conclusion:

      Overall, this reviewer considers that the study requires a major revision, including improvements in numerical methods, modeling choices, and checks for physiological behavior. Nevertheless, the provided tables with averaged values from the UK Biobank and the presented validation strategy could be valuable to the research community.

      Reviewer #2 (Public review):

      The authors present an interesting study on calibrating and validating a biventricular cardiac electromechanical model. This is an important contribution, but some questions remain about the quantitative validation and verification aspects of the study.

      Major comments:

      (1) The title and paper stress the importance of validation on several occasions. However, the actual validation performed is limited to the section in lines 427-439. Furthermore, it is entirely qualitative, making assessing the model's quality difficult. Most of the paper is focused on sensitivity analysis, which is also interesting but unrelated to validation. Can you include a quantitative comparison with deformation biomarkers? E.g., spatially quantify strain differences between simulation and in vivo data, or overlay the current configuration of the geometry with MRI in various views, and calculate a displacement error norm.

      We thank the reviewer for this comment.

      We have strengthened the quantitative aspect of the validation by reporting the peak simulated strain values for each component and comparing them against the physiological ranges compiled in Table 2. Specifically, the simulated peak strains were: E_ff ≈ -0.20, E_cc ≈ -0.15, E_rr ≈ +0.15, and E_ll ≈ -0.23. These show broad agreement with the in vivo reference ranges from Moulin et al. (2021), noting that the reference ranges are derived from a cohort of 30 subjects and therefore represent a relatively narrow population sample. Shortening strains (fibre and circumferential) are in good agreement, while radial strain is underestimated. We have noted this as a limitation. We have also indicated that a further validation would include a fully quantitative spatial comparison, such as a displacement error norm or voxel-wise strain difference map. This would require access to the raw image data and patient-specific geometry registration, which is beyond the scope of the current study.

      (2) You mention the ASME V&V40 standards throughout your paper. Yet, you only address the "second V" validation, ignoring the "first V" verification. How did you ensure that your computational models are implemented correctly?

      Thank you for raising this point. We have now included a section on model verification to the manuscript at where we perform benchmarking simulation using the Land (2015) passive inflation benchmark. We also provide a mesh subdivision analysis of the final calibrated model. Additional verifications and previous sensitivity analyses using the same numerical scheme with idealised ellipsoid geometries are also referenced in the verification section, to provide additionally confidence.

      (3) All parameters discussed in this publication are physical parameters. What is the sensitivity of your model outputs concerning computational parameters?

      Numerical analyses for the Alya solver used in this study has previously been published in works including Levrero et al (2021), which performed sensitivity analyses in a truncated ellipsoid geometry, and Santiago et al (2018), which demonstrated mesh convergence in a cantilever. We have updated the manuscript to point the reader to these studies.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major concerns:

      (1) Active stress scaling:

      The initial value for T_ref appears to be 120kPa * 10, which would be ten times the literature value fitted to human contraction data. Additionally, Table 2 lists a range of [1200-2400], which is 10 to 20 times the literature value.

      This discrepancy suggests that other model parameters, model assumptions, or the numerical scheme may be inadequate. In contrast, similar calibrations using comparable models (ToRORd-Land) in other works, such as Strocchi et al. [29], yielded T_ref values close to the literature value.

      We thank the reviewer for this comment. As discussed in our response to the public review comment 3a, the elevated T_ref scaling warrants explanation.

      We note that Strocchi et al. use a four-chamber geometry include atrial mechanics and a different pericardial constraint, any of which could contribute to differences in the required T_ref scaling. The elevated scaling in our model likely reflects a combination of factors including the absence of poro-elastic behaviour, simplifications in pericardial constraint, and the lack of atrial mechanics, rather than volumetric locking alone. We have added text to the discussion acknowledging this more explicitly and have flagged planned additional benchmarking of the dynamic orthotropic scheme as future work.

      (2) Non-physiological results, see Figure 1:

      In a healthy heart, RV stroke volume should approx. match LV stroke volume. This is clearly not the case in Figure 1B, where the RV EF is also notably low at 35%.

      Consequently, the study fails to reproduce healthy pressure-volume loops, undermining its claim to create a credible cardiac electromechanical digital twin. Hence, also the "Question of interest" posed in line 206 must be answered with a clear "No".

      Matching stroke volumes should be a primary calibration goal.

      We thank the reviewer for this comment. We agree that stroke volume balance is an important physiological criterion, and we have added it explicitly to the framework criteria in the updated manuscript, noting that our current model evaluation does not satisfy it. This is precisely the kind of transparent appraisal the V&V40 framework is designed to produce: a systematic accounting of which criteria are met and which require further development. A framework that only gets applied to models that pass all criteria would be selection-biased and less informative to the community.

      However, we respectfully disagree that the question of interest must be answered with a clear 'No'. We draw the reviewer's attention to the quantities of interest defined in the paper, which are predominantly left ventricular biomarkers, reflecting the intended scope of the calibration framework. The framework successfully reproduces these defined quantities of interest, and the LV pressure-volume loops, strain, volumes and ejection fraction are all within physiological ranges and well-matched to reference data. These quantities of interest were selected based on their clinical implications in cardiac diseases, as detailed in Table 3.

      We agree that stroke volume balance is an important physiological requirement for a fully calibrated biventricular model, and we have strengthened the future work and limitations section accordingly.

      (3) Inadequate numerical framework:

      (a) Monodomain model: The geometries from Rodero et al. [25] have an average edge length of 1mm. It is known that such a coarse resolution leads to inaccurate EP results. It is not mentioned if the authors refined that geometry to an appropriate resolution or used an Eikonal model to mitigate this issue.

      As explained earlier, our ECG simulations are robust against coarse mesh resolutions since we use an Eikonal solution to prescribe the activation times on the endocardial surface, as performed in Camps et al. (2024), and we tune the diffusivity parameters in the model such that the correct conduction velocities are reached. We have added a figure in the appendix of this manuscript to show that by increasing the mesh resolution by one subdivision, we get virtually identical ECG simulations. While this mesh resolution may not be sufficient for simulations of more complex behaviour, such as re-entry and fibrillation patterns, it is sufficient for this study. We have noted this in the methods section.

      (b) Material law:

      - recent publications show that an unsplit deformation gradient for the anisotropic contribution is beneficial to reduce locking effects, see, e.g., Gueltekin et al. Computational Mechanics 63, no. 3 (2019): 443-53. https://doi.org/10.1007/s00466-018-1602-9.

      - K_ct is a penalty parameter to enforce some degree of incompressibility. Results are highly dependent on the grid size and the finite element formulation due to locking effects.

      As the authors write: "In our simulations, we saw that the LVEF was strongly sensitive to changes in the incompressibility of the tissue (Kct), such that an increase in compressibility of the myocardial tissue helped to increase LVEF." Which exactly points to the issue of locking effects.

      So an option would be to use a finer grid or a more adequate numerical scheme with quadratic finite elements, as eg. in [5] Fedele et al., or [6] Gerach et al,. or stabilized elements as in Karabelas et al. CMAME 394 (2022) https://doi.org/10.1016/j.cma.2022.114887.

      Overall, this does not point to "limitations in using ex vivo tissue measurements to represent in vivo function" but to limitations in the numerical setup. In fact, with an adequate numerical scheme, the simulations should be largely insensitive to the choice of this penalty parameter K_ct. See, e.g., Karabelas et al. above, where the authors varied K_ct from 650kPa to infinity (representing an incompressible material), and there is no visible influence on the PV loops.

      We investigated this point using the Alya solver, and we found that the mesh resolution did not alter the LVEF, and our benchmark simulations against Land (2015) did not show the existence of the volumetric locking issue that the reviewer refers to. It is possible, however, that such an effect exists in the elastodynamic orthotropic framework but not in the incompressible and transversely isotropic framework that the Land (2015) benchmarks were set up in. Future analyses could focus on performing additional benchmarking against more recent elastodynamic benchmarks, such as presented in Arostica (2025). We have updated the limitations text in our manuscript to reflect this and to cite relevant literature on this issue.

      (4) Boundary conditions:

      "This was a simplified version of the method [28], which uses an exponential decay formulation at the 'edge' of the pericardial constraint rather than a step function": I don't really see this in the cited work [28] which gives a spatially varying Robin-type boundary condition at the whole epicardium (i.e. regional scaling of normal springs stiffness based on image-derived motion from CT images) and not only at the edge.

      This is motivated by the fact that the pericardial tissue is in contact with various organs of different material properties. Not using spatially varying pericardial parameters is a limitation that might lead to non-physiological deformations, see also Pfaller et al. Biomechanics and Modeling in Mechanobiology 18 (2019): 503-29. https://doi.org/10.1007/s10237-018-1098-4.

      We thank the reviewer for this point and we have corrected the manuscript accordingly. To clarify: our implementation applies a uniform Robin spring constraint along the majority of the epicardial surface with zero constraint at the base, which is conceptually similar to Strocchi et al. [28]. The key difference is that Strocchi et al. use a smooth gradient transition from uniform constraint to zero constraint near the base, whereas our implementation uses an abrupt step transition. We acknowledge that a smooth spatially varying transition would more accurately represent the frictionless pericardial contact and have noted this as a limitation in the manuscript with reference to Pfaller et al. [41].

      Also check:

      - line 117: Gamma_valve_epi is introduced but not used. Was there any boundary condition defined on this valve plane?

      - the third equation, maybe (0,T] missing.

      - line 120: epicardium instead of endocardium.

      These errors have been corrected in the updated manuscript. No boundary conditions were applied on the epicardial surface of the valve plugs, the reference to gamma_valve_epi has been removed.

      (5) Reference geometry:

      The choice to scale the mesh to a lower volume for the unloading procedure seems questionable. This approach does not ensure that the reloaded mesh aligns with the mesh derived from image data. As a result, the geometry used for the simulations is no longer truly patient-specific.

      This mismatch is a significant limitation, as there are established methods available to achieve a proper unloaded configuration, as, e.g., in

      Marx et al. Journal of Computational Physics 463 (2022): 111266. https://doi.org/10.1016/j.jcp.2022.111266, and

      Regazzoni et al. Journal of Computational Physics 457 (2022): 111083. https://doi.org/10.1016/j.jcp.2022.111083.

      As our study aimed at creating a framework for calibration and validation in data-scarce scenarios such as it is often the case in the clinical context, using a compilation of multi-modal data from difference sources, rather than a specific method of personalisation, we did not feel it appropriate to invest significant energy to identify a patient-specific resting geometry, but rather felt that it was important for the resting geometry to fall within population values in terms of diastasis volume. We have clarified this issue in the manuscript and softened claims to Digital Twins in this study. The limitation has been addressed in the updated manuscript, and future work could further address this point.

      (6) Calibration procedure:

      There are apparent flaws in the calibration procedure, or it is not described in sufficient detail.

      We thank the reviewer for raising this point and we have substantially revised the calibration description in the manuscript to clarify the rationale behind each step.

      (a) Step 1: "Sample..." Why? kws and Cal50 are not the most significant parameters in the sensitivity analysis. Kct is a penalty parameter dependent on the numerical framework as described above; "ejection pressure threshold" was never mentioned, is it "P ejection LV" in Table 2? Aiming just for the highest LVEF might neglect non-physiological responses to parameter changes.

      While kws and Cal50 are not the single most significant parameters for LVEF in isolation, they were grouped in Step 1 because they affect both LVEF and peak systolic pressure simultaneously through cross-bridge cycling rate and residual active tension, making it necessary to sample them jointly rather than sequentially. Kct was included because myocardial compressibility affects wall thickening and therefore stroke volume. The ejection pressure threshold is P_ejection_LV in Table 1 and has now been described explicitly in the methods section. Regarding the concern about non-physiological responses: the action potential duration and active tension were monitored throughout calibration and verified to remain within physiological ranges, as now noted in the manuscript.

      (b) Step 2: As systolic pressure is directly dependent on arterial resistance for a 2-element Windkessel model, a uniform sampling approach might not be the best choice here.

      We acknowledge that uniform sampling may not be the most efficient approach for Step 2. However, since arterial resistance influences not only peak systolic pressure but also stroke volume and therefore LVEF, a more targeted approach focusing solely on pressure matching could compromise the LVEF achieved in previous steps. Uniform sampling allowed us to select the value that best balanced both quantities simultaneously.

      (c) Step 3: The authors mention in line 527: "A four-fold increase in GCaL caused an eight-fold increase in cellular active tension peak". An increase in active tension peak results in higher LVEF. So this step is likely to yield the upper boundary of the GCaL interval.

      The reviewer is correct that Step 3 tends to yield a high GCaL value. This was intentional — GCaL was used as a last resort to achieve physiological LVEF after Steps 1 and 2, since the model consistently undershot the target. The upper boundary of the sampled GCaL interval corresponds to a two-fold increase, which remains within the physiological variability bounds applied in previous studies. The resulting action potential duration was verified to remain within physiological ranges.

      (d) Step 4: Why again k_ws? It is not the most significant parameter in the SA.

      kws was resampled in Step 4 not to increase LVEF further, but to specifically target peak ejection rate and dP/dtmax, which were not adequately matched after Step 3. kws is the dominant parameter affecting these ejection dynamics biomarkers in the sensitivity analysis. Resampling at this stage allowed fine-tuning of ejection dynamics while maintaining the LVEF achieved in previous steps.

      (e) Step 5: As far as I can tell, the "diastolic volume change parameter" was mentioned the first time here.

      The diastolic volume change parameter C_pLAV has now been described in the methods section in the Phase 5 passive filling description, where it appears as the inverse of the penalty term controlling the rate of return to diastasis volume in the left ventricle.

      The whole calibration procedure seems to aim for the highest LVEF, and final values of the calibration parameters are not given.

      We note that the calibration procedure does not aim solely for the highest LVEF. As described above, the sequential strategy targets multiple quantities of interest in order of clinical importance: LVEF, peak systolic pressure, peak ejection rate, and peak filling rate, with each step designed to improve a specific subset of biomarkers without compromising those already matched. The final calibrated parameter values are reported in Figure 1F of the revised manuscript.

      (7) Novel features in this paper are actually scarce. A way more advanced calibration strategy with a whole heart model, emulators, and also the ToRORd-Land model was already presented in the study by Strocchi et al. [29]. The calibration to ECGs was presented by some of the same authors in Camps et al. [15], and the analysis of cellular effects was already published in several studies by the same group and in other publications, e.g., by the groups of Severi et al.

      The systematic compilation of credibility criteria spanning ECG morphology, pressure-volume characteristics, strain and displacement represents a novel contribution in itself, providing the field with a reusable evaluation framework. Furthermore, the present study is designed to yield mechanistic insight into how parameters at different scales influence both electrical and mechanical outputs simultaneously. This goal was not tackled in previous publications, which covered individual components, including ECG calibration in Camps et al. [15] and global sensitivity analysis with whole-heart models in Strocchi et al. [29] with no ECG consideration.

      Thus, the work by Camps et al. on ECG calibration was purely electrophysiological and did not investigate the influence of mechanical or haemodynamic parameters on ECG morphology in a fully coupled electromechanical framework. While the effect of mechanical parameters on ECG has been explored by others (e.g. Favino, 2016), this has not previously been examined alongside the relative importance of cellular, mechanical and haemodynamic parameters on pressure-volume characteristics within a single coupled framework. While Strocchi et al. present an emulation strategy, they did not address ECG biomarkers. This distinction is now stated explicitly in the introduction, where we position the present study relative to Camps et al. and Strocchi et al.

      (8) How could the calcium sensitivity Cal50 have such a drastic effect on diastolic function, i.e., filling and end-diastolic volume? As far as I understand from the description, the simulation starts with Phase 0 (loading), Phase 1 (atrial filling), and then in Phase 2, electrical activation ensues and active contraction develops, see also the section starting in line 165. Based on this description, I would expect the end-diastolic volumes to be identical across all Cal50 values. Or are the PV loops shown actually limit cycles established over simulations with multiple beats? This point wasn't explicitly clarified in the manuscript.

      Calcium sensitivity (Cal50) affects not only systolic active tension development but also diastolic residual active tension, i.e. the degree to which the muscle remains partially activated at end diastole. Higher Cal50 values increase this residual tone, effectively stiffening the myocardium during diastolic filling and reducing end-diastolic volume. This mechanism is well established as a contributor to diastolic dysfunction in heart failure [88]. We have clarified this in the manuscript and also clarified that the PV loops shown are single-beat simulations, not limit cycles, with the end-diastolic volume determined by the prescribed filling pressure alongside the passive and residual active stiffness of the myocardium.

      Minor concerns:

      (9) Line 29: The values provided: LVEF of 51%, EDV of 110 mL, and ESV of 50 mL are inconsistent. If these values are all related to the LV, the calculated LVEF should be approximately 54.55%, not 51%.

      The values quoted in the original abstract were rounded approximation, this has been corrected to report EDV=105 mL and ESV=51 mL, which are consistent with the simulated LVEF of 51%.

      (10) "Electromechanical cardiac Digital Twins have had broad applicability ..."

      Many of the cited works here are not true "Digital Twins" but rather static, non-patient-specific models of cardiac electromechanics. In some cases, the geometry may be derived from patient data, but this alone does not qualify the model as a digital twin.

      This sentence in the introduction has been rephrased as ‘Electromechanical cardiac models have had broad applicability...’. Furthermore, as stated earlier, we have removed explicit claims of Digital Twin from the paper while retaining the fact that this study provides a significant step towards rigorous credibility assessment of the high-fidelity electromechanical models that make Digital Twin construction possible.

      (11) While in the abstract and the conclusion, the authors mention "uncertainty quantification", it is mostly a sensitivity analysis that was performed in the paper.

      We have updated the text to say ‘sensitivity analysis’ where appropriate in the abstract, results, and conclusion, and replaced ‘uncertainty ranges’ with ‘variability ranges’ throughout. However, since the sensitivity analyses were performed over biologically informed ranges derived from population variability in the literature, the results are informative about how uncertainty in model inputs propagates to uncertainty in simulated biomarkers. We have therefore retained the framing of sensitivity analysis as a first step towards uncertainty quantification in the abstract and conclusion, and have added a clarifying sentence to the methods to this effect.

      (12) Line 98: As far as I can tell, the conduction velocity assigned to the endocardial surface - intended to mimic the Purkinje fiber network - is never specified. In the section beginning at line 294, only the transmural conduction velocities are reported.

      The endocardial conduction velocity has been specified in the methods section: Purkinje-myocardial junctions were modelled using a fast endocardial activation layer with isotropic conduction velocity of 300 cm/s.

      (13) Line 198, Table1:

      (a) "21/02/2025 11:09:00 AM" on two occasions is maybe not intended

      This has been removed.

      (b) For easing up comparisons, units should be consistent between the initial value and the literature ranges, e.g., PV control parameters, heart rate.

      Units have been made consistent between the initial values and literature ranges throughout Table 1.

      (14) Line 232: "... have already been used to calibrate and validation ...".

      This has been corrected.

      (15) Line 279, Table 2: This table of variability ranges is not entirely clear and could be improved:

      Table 2 has been combined with Table 1 such that the variability ranges sit next to the literature values, for ease of comparison.

      (a) "21/02/2025 11:09:00 AM" is maybe not intended.

      This has been removed.

      (b) use of units should be improved; sometimes it's given in the first column, sometimes in the second column (arterial resistance, compliance), then for k_epi it should be either kPa or kPa/cm.

      Units have been made consistent and the units for k_epi has been added in Table 1.

      (c) units should also be consistent throughout the paper, e.g. in Figure 1 E arterial resistance is Barye.ms/mL while in Table 2 it is mmHg.ms/mL.

      Barye has been removed and replaced by corresponding kPa values throughout the manuscript. This was in the original manuscript since the Alya simulation software were in units of cm, s, g, Barye.

      (d) it is also not clear how variability ranges were chosen; e.g., for arterial compliance,e literature ranges are 0.2-2.73 while the chosen range is [0.1,0.2].

      The previous ranges were chosen to achieve better LVEF. We have now updated the variability ranges to be purely based on literature values and updated the sensitivity analysis results. The ranges are now presented in Table 1 alongside the literature values for ease of comparison.

      (e) For Kct, the initial value in Table 1 is 5000kPa, the literature values are between 10 and 3333, and then the variability range is [10,500]? I guess there is a typo in one of these values.

      This has been corrected in the new Table 1.

      (f) Table 1 and 2 are in parts redundant.

      Table 1 and 2 have been combined into a single new Table 1.

      (15) Line 290: It should be uvc_l for the longitudinal coordinate.

      This has been corrected.

      (16) Line 388, Table 3, regarding values for pressure volume from reference [49]:

      (a) the number of participants is 800, including males and females; not only females, see also Table 12 https://jcmr-online.biomedcentral.com/articles/10.1186/s12968-017-0327-9/tables/12

      (b) why using female values here while having mixed sex for most of the others? Because the model is female?

      The reviewer is correct that reference [49] reports values from a mixed-sex cohort of approximately 800 participants. We used the female-specific values from Table 12 of that reference because the biventricular mesh used in this study was derived from a female subject, making sex-matched reference values the most appropriate comparison. This has been clarified in the manuscript.

      (17) Figure 5: What is Jup; why did you choose 0.93 x Jup as reference? Also in Figure 4, why did you use 0.93 x GCal as a reference?

      J_up refers to the SERCA<sup2+</sup> reuptake current, which has been relabelled as SERCA throughout the manuscript for consistency. The reference value of 0.93× was used because the sensitivity analysis sampled parameters uniformly between 50% and 200% of baseline using a fixed number of samples, and no sample fell exactly at 1.0×. The closest sampled value was 0.93×, which was therefore used as the reference. This has been clarified in the figure caption.

      (18) Line 377: The link to the GitHub repository does not work.

      This link has now been made publicly available.

      (19) Line 397, Table 3: for the sake of completeness, all abbreviations should be included: e.g., SVL, ESP, EDV, ESV are not included.

      This has been written out in full in the new Table 2.

      (20) Tick marks in many figures are not readable, e.g., Figure 5 and all the Figures in the appendix.

      Tick mark sizes and line widths have been increased across Figures 4, 5, and all appendix figures. The figures have been replotted and updated in the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      Minor Comments:

      (1) The provided GitHub link https://github.com/jennyhelyanwe/Alya_input_setup/ does not work, potentially because the repository is private. It would be nice to see the repository during the review.

      This link has now been made publicly available.

      (2) Table 3: Can you include the simulation outputs obtained for validation (with an error indication)? This would summarize the validation that's currently spread out over the results section.

      A new Table 3 has been added to the manuscript under the validation section, summarising the simulated values for all deformation and strain biomarkers alongside their reference ranges. The calibration and validate datasets are now reported separately in Tables 2 and 3, respectively.

      (3) Figure 1: Add axis labels to all plots.

      Axis labels have been added to all subplots in Figure 1 in the revised manuscript. Simulated pseudo-ECG amplitudes are normalised and therefore dimensionless.

      (4) Figure 2: Simulated and in vivo strains with exactly the same axes (size, range, ticks) and add grid lines to enable a comparison. Add the mean values of each in the other plot.

      The revised Figure 2 now includes the median in vivo strain values from Moulin et al. overlaid as a red dashed reference line on the simulation panels, enabling direct visual comparison. The simulated mean could not be overlaid on the in vivo panels as the original Moulin et al. figure data are not publicly available for replotting. Exact axis matching was not applied as this would cause some simulated curves to fall outside the visible range, obscuring the model behaviour.

      (5) Figure 3: The thickness (relative importance of the connections) is impossible to see in this plot. Instead of having gray background connections, remove them entirely below a certain threshold. Make the differences in thickness more pronounced or introduce a continuous color scale for the magnitude of the positive or negative correlation. Alternatively, you could rank the parameters from least to most important in each subfigure A-D and/or provide some numeric values.

      Figure 3 has been updated. All non-significant connections (|r| < 0.6 or p > 0.05) have been removed entirely, and gray lines have been removed in each subfigure, making the significant relationships clearer. A continuous blue-to-red colour scale has been applied to indicate the direction of correlation (blue: negative, red: positive), with line thickness proportional to the magnitude of the r-value.

      (6) Figure 3 and Table 2: Why were material parameters b, bf, bs, and bfs omitted from this study (but included a, af, as, and afs)?

      The b parameters (b, bf, bs, bfs) appear in the exponent of the Holzapfel-Ogden constitutive law and are strongly coupled to the a parameters (a, af, as, afs), which carry units of kPa. In practice, the b parameters can only be reliably identified from ex vivo multiaxial stretch experiments, whereas the a parameters can be estimated from clinical imaging data. Since our study focuses on calibration and validation in a clinical data setting, we included only the a parameters in the sensitivity analysis, consistent with previous personalisation studies.

      (7) Figure 4: What do the dotted lines represent?

      The dotted lines in Figure 4E highlight the increased longitudinal shortening with increasing GCaL, showing the basal plane moving towards the apex while the apical position remains unchanged due to the pericardial constraint. This has been clarified in the figure caption.

      (8) Figures 4, 5, A2-45: Can you use a continuous color scale (e.g., from blue to red) for low to high parameter uncertainty?

      A continuous blue-to-red colour scale has been applied to Figures 4, 5, and all appendix figures A2–A6, where blue indicates the lowest parameter value and red indicates the highest. A colour bar has been added to each figure for reference.

    1. eLife Assessment

      The study presents important findings from a very rich EEG-fMRI dataset including 107 participants, which was collected during nocturnal naps. Using overall solid methods, the authors link activity in memory related brain regions (e.g., hippocampus, thalamus, and medial prefrontal cortex), as well as their functional connectivity, to the occurrence of canonical sleep rhythms (spindles and slow oscillations) during non-rapid eye movement (NREM) sleep. This work will be of broad interest to researchers studying sleep, memory, and related domains.

    2. Reviewer #1 (Public review):

      Wang et al., recorded concurrent EEG-fMRI in 107 participants during nocturnal NREM sleep to investigate brain activity and connectivity related to slow oscillations (SO), sleep spindles, and in particular their co-occurrence. The authors found SO-spindle coupling to be correlated with increased thalamic and hippocampal activity, and with increased functional connectivity from the hippocampus to the thalamus and from the thalamus to the neocortex, especially the medial prefrontal cortex (mPFC). They concluded the brain-wide activation pattern to resemble episodic memory processing, but to be dissociated from task-related processing and suggest that the thalamus plays a crucial role in coordinating the hippocampal-cortical dialogue during sleep.

      The paper offers an impressively large and highly valuable dataset that provides the opportunity for gaining important new insights into the network substrate involved in SOs, spindles, and their coupling.

      Comments on latest version:

      The authors have substantially revised their manuscript and sufficiently addressed all of my previous concerns. I have no further comments.

    3. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Wang et al., recorded concurrent EEG-fMRI in 107 participants during nocturnal NREM sleep to investigate brain activity and connectivity related to slow oscillations (SO), sleep spindles, and in particular their co-occurrence. The authors found SO-spindle coupling to be correlated with increased thalamic and hippocampal activity, and with increased functional connectivity from the hippocampus to the thalamus and from the thalamus to the neocortex, especially the medial prefrontal cortex (mPFC). They concluded the brain-wide activation pattern to resemble episodic memory processing, but to be dissociated from task-related processing and suggest that the thalamus plays a crucial role in coordinating the hippocampal-cortical dialogue during sleep.

      The paper offers an impressively large and highly valuable dataset that provides the opportunity for gaining important new insights into the network substrate involved in SOs, spindles, and their coupling.

      Thank you for this encouraging assessment. We appreciate your recognition of the value of the dataset and of the questions it allows us to address. Below, we respond to each of your points directly and revise the manuscript accordingly.

      Comments on revisions:

      Re 1: The revised introduction now cites a couple of papers but discusses them only very superficially, lumping together several studies with very different key results. This is still not very informative for the reader and does not sufficiently acknowledge previously published work. Here are two examples to illustrate this:

      (a) "These studies have generally reported that slow oscillations are associated with widespread cortical and subcortical BOLD changes, whereas spindles elicit activation in the thalamus, as well as in several cortical and paralimbic regions." Several studies even showed e.g., a clear activation of the hippocampus and parahippocampal gyrus associated with spindles, not just the thalamus

      Thank you for this comment. We agree that our previous sentence was too broad and did not sufficiently reflect the range of findings in the sleep literature. We have therefore rewritten the Introduction to state explicitly that spindle-related BOLD changes have been reported not only in the thalamus, but also in cortical and paralimbic regions, including the hippocampus and parahippocampal gyrus (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Schabus et al., 2007).

      Introduction, Page 3-4, Lines 58-62

      “Consistent with this view, prior human EEG-fMRI studies have reported spindle-related activation not only in the thalamus, but also in the hippocampus and adjacent parahippocampal gyrus (Bergmann et al., 2012; Schabus et al., 2007). Spindle-related activity has also been linked to striatal engagement, suggesting a broader network that may support memory-related processing during sleep (Fogel et al., 2017).”

      Introduction, Page 4, Lines 71-78

      “Previous EEG-fMRI studies on sleep have examined both global sleep characteristics (Hale et al., 2016; Moehlman et al., 2019) and the neural correlates of specific waves, including slow oscillations and spindles. These studies have generally shown that slow oscillations are associated with widespread cortical and subcortical BOLD changes (Czisch et al., 2009; Ilhan-Bayrakcı et al., 2022; Picchioni et al., 2011), whereas spindles have been linked not only to thalamic activation but also to cortical and paralimbic regions, including the hippocampus and parahippocampal gyrus (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Schabus et al., 2007).”

      Introduction, Page 5, Lines 103-106

      “This coupling was associated with increased activation in both the thalamus and hippocampus, with functional connectivity patterns suggesting thalamic coordination of hippocampal-cortical communication, in line with prior EEG-fMRI studies of spindle-related activity (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Schabus et al., 2007).”

      (b) "Although these findings provide valuable insights into the BOLD correlates of sleep rhythms, they often do not employ sophisticated temporal modeling (Huang et al., 2024) [, ...]." - previous studies have used e.g., spindle event-related regressors with individual spindle amplitudes as parametric modulators, first and second order derivatives of the HRF function, as well as PPI connectivity analyses, which I would consider rather sophisticated temporal modelling.

      We agree that several previous studies have already employed sophisticated modelling approaches, including parametric modulation, HRF derivatives, and PPI analyses (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Picchioni et al., 2011). Our intention was not to suggest that such methods are absent from the literature.

      Rather, we aimed to highlight that most prior work has focused on modelling individual SO or spindle events, whereas explicit modelling of their temporal interaction (e.g., SO-spindle coupling) has been less commonly addressed. We have revised the sentence to clarify this point more precisely.

      Introduction, Page 4 Lines 78-82

      “Although these findings provide important insight into the BOLD correlates of sleep rhythms, most previous studies have focused on individual oscillatory events rather than explicitly modelling their temporal interaction (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Picchioni et al., 2011). Only a few recent studies have begun to examine coupling between rhythms directly, for example Huang et al. (2024).”

      Re 4+9: The short overall recordings in some subjects on the one hand and the large number of spindles and SOs detected in N1 sleep stages are still highly concerning, in fact even more so, now that the actual numbers have been provided in the Supplementary Tables. Either the sleep staging or the detection of SO and spindle events must be incorrect. I understand that for specific EEG analysis and fMRI modelling purposes sometimes slightly different thresholds are used as compared to clinical sleep staging, but several parameters here are alarmingly off.

      (a) Given that proper NREM sleep (N2+N3) is the relevant stage for the analyses conducted in this paper, some of the N2+N3 durations are very short (eg 7-8 min) while those subjects' results have the same impact on the group level analyses as those with >100 min of N2+N3. Either subjects with very little relevant data (not overall recording time but N2+N3 time) should be excluded or weighting subject data for the group analyses according to the amount od contributed data should be done.

      Thank you for the suggestion. It is true that participants with very little N2/3 sleep could contribute noisier subject-level estimates to the group analysis. We therefore checked this directly. Only three participants contributed less than 10 min of N2/3 sleep, and excluding them did not change the main results. For example, hippocampal activation during SO-spindle coupling remained significant after exclusion, t<sub>(103)</sub> =2.50, p = 0.0071, compared with t<sub>(106)</sub> = 2.50, p = 0.0070 in the full sample. We have added this control analysis to the Results so that the robustness of the group findings is explicit in the manuscript.

      Results, Page 11-12, Lines 238-250

      To ensure the results were not driven by individual differences or parameter selection, we conducted a series of control analyses. First, we excluded participants with less than 10 minutes of N2/3 sleep. Only three participants met this criterion, and their exclusion did not change the main results. For example, hippocampal activation during SO-spindle coupling remained significant (t<sub>(103)</sub> = 2.50, p = 0.0071), comparable to the full sample (t<sub>(106)</sub> = 2.50, p = 0.0070). Second, because the absolute number of detected SO-spindle coupling events depends on the SO detection threshold, we examined whether the main EEG-fMRI results were sensitive to this parameter. To this end, we varied the SO percentile threshold and reconstructed the EEG-informed GLM at each level. Hippocampal activation during SO-spindle coupling remained significant across a range of thresholds (71st-80th percentile; Fig. S6). Third, to test whether the results depended on the use of a single lateralised frontal electrode, we repeated the EEG-informed fMRI GLM using events detected from Fz. Hippocampal activation during SO-spindle coupling again remained significant (t<sub>(106)</sub> = 2.47, p = 0.0076), closely matching the original F3-based result (t<sub>(106)</sub> = 2.50, p = 0.0070).

      (b) The authors argue that the SO and spindle detection algorithms are valid since widely used and that they were developed for N2+N3 stages, which is why they will also detect events in other stages: "While, because the detection methods for SO and spindle are based on percentiles, this method will always detect a certain number of events when used for other stages (N1 and REM) sleep data, but the differences between these events and those detected in stage N23 remain unclear." I do agree that with very liberal thresholds, also SO and spindle vents may be detected in other stages, but it shouldn't be that many. If the percentiles of amplitude thresholds were defined based on properly scored N2+N3 stages only, very few events should be detected (erroneously!) in N1, as the occurrence of K-complexes (isolated SOs) and spindles per definition makes it N2, and during REM sleep only very few spindles and SOs are allowed to occur, without scoring it NREM instead. For the first subject (just as example, but with similar numbers for the rest of the sample), reveals as many as 60 SOs and 31 spindles within 8 min of N1 sleep (Table S2) as well as 13 SOs and 7 spindles within 2 min of REM sleep (Table S4). These numbers are completely unrealistic and question the correctness of the sleep staging as well as the physiological relevance of the EEG graphoelements identified as SO and spindles. It also completely undermines the interpretability of the respective event regressors for the fMRI analyses.

      (c) Likely, given the large numbers of coupled SO-spindle events and the apparently very low amplitude criteria for event identification, also the number of SO-spindle couplings is likely severely overestimated.

      We thank the reviewer for raising this important point. We agree with you that the original stage-wise percentile thresholding could inflate the apparent number of SOs and spindles outside N2/3 sleep. In the original analysis, the thresholds were estimated separately within each sleep stage. As you point out, this procedure can force the detector to label a relatively large number of events in N1 and REM, even when those waveforms should not be interpreted as canonical N2/3 SOs or spindles. We have therefore revised the detection procedure. Following your concern and Reviewer 2’s suggestion, the SO and spindle thresholds are now defined only from N2/3 sleep within each participant, where SOs and spindles are most abundant and physiologically expected to occur. These fixed N2/3-derived thresholds were then applied unchanged to N1 and REM for descriptive reporting. This avoids the artificial normalisation of event detection across sleep stages that can arise when each stage has its own percentile threshold. And we have revised all relevant sections of the manuscript, including “[Results, Page 6-7 Lines 134-148]; [Fig. 1e]; [Results, Page 9 Lines 175-191]; [Fig. 2b]; [Methods, Page 25-27, Lines 567-604]; [Fig. S2-S4]; [Table S2, S4].”

      With this revised procedure, detections outside N2/3 are clearly lower than those in N2/3. The mean densities are 2.95 SOs/min, 2.71 spindles/min, and 0.75 coupling events/min in N1, and 2.07 SOs/min, 1.81 spindles/min, and 0.43 coupling events/min in REM. We agree with you that the remaining detections in N1 and REM should not be treated as physiological equivalents of canonical N2/3 SOs, spindles, or SO-spindle complexes. We therefore report them only as descriptive detector outputs obtained under a fixed N2/3-derived threshold (see Table S2, S4 in the revised manuscript). We do not use them to support any physiological claim about SO-spindle coupling in N1 or REM.

      This point is also important for the fMRI analyses. You are right that inflated N1 or REM detections would undermine the interpretability of event regressors if those detections entered the EEG-informed fMRI models. They did not. All EEG-informed fMRI GLM and PPI analyses were restricted to N2/3 sleep, where SOs, spindles, and their coupling are physiologically expected and where the detection thresholds were defined. Thus, the central fMRI event regressors were based only on N2/3 events, not on detections from N1 or REM.

      We also agree with you that the absolute number of detected SO-spindle couplings depends on the chosen detection threshold. For this reason, we tested whether the main EEG-fMRI result depended on the specific detector setting. Hippocampal activation during SO-spindle coupling remained significant when the SO detection threshold was varied between the 71st and 80th percentiles, as shown in Fig. S6. We therefore do not argue that the detector provides a uniquely correct absolute count of SOs, spindles, or coupling events in every sleep stage. Our conclusion is more specific. The main N2/3 EEG-fMRI finding is robust across a reasonable range of SO detection thresholds, detections in N1 and REM are reported only descriptively, and the physiological interpretation of SO-spindle coupling is restricted to N2/3 sleep.

      Results, Page 6-7 Lines 134-148

      “Each sleep stage is characterised by distinct spectral properties and rhythmic waveforms, serving as physiological markers (Fig. 1c). Because SO and spindle detection relies on amplitude-based percentile thresholds, we avoided estimating separate thresholds within each sleep stage. Instead, for each participant, the SO and spindle thresholds were defined from N2/3 sleep only, where these rhythms are most abundant and physiologically expected, and the same fixed thresholds were then applied to N1 and REM for descriptive comparison.”

      “Under this fixed N2/3-derived thresholding, detected SOs and spindles were larger and more frequent in N2/3 than in N1 or REM. SO and spindle amplitudes were significantly higher during N2/3 sleep (SO: 25.59 ± 1.49 μV; spindle: 7.39 ± 0.27 μV) than during N1 (SO: 20.15 ± 2.32 μV; spindle: 5.23 ± 0.27 μV) and REM sleep (SO: 19.84 ± 1.22 μV; spindle: 5.60 ± 0.22 μV; all p < 1e-4; Fig. 1e, Fig. S2). The corresponding event densities showed the same pattern, with 9.64 ± 0.25 SOs/min and 4.19 ± 0.10 spindles/min in N2/3, compared with 2.95 ± 0.16 SOs/min and 2.71 ± 0.14 spindles/min in N1, and 2.07 ± 0.17 SOs/min and 1.81 ± 0.14 spindles/min in REM (all p < 1e-4). We therefore report detections in N1 and REM only as descriptive outputs of the detector under a fixed N2/3-derived criterion, rather than as physiological equivalents of canonical N2/3 SOs or spindles.”

      Fig. 1 legend, Page 8, Line 166-172

      “e, Amplitudes (μV) of detected SOs (left) and spindles (right) across sleep stages. SO and spindle detection thresholds were defined from N2/3 sleep within each participant and then applied unchanged to N1 and REM for descriptive comparison. Detections in N1 and REM should therefore be interpreted as detector outputs under this fixed N2/3-derived criterion. The SO amplitudes were measured from the 0.16-1.25 Hz filtered EEG data, and spindle amplitudes were measured from the 12-16 Hz filtered EEG data. Each dot represents an individual participant. Error bars indicate SEM. *** p < 0.001.”

      Results, Page 9 Lines 175-191

      “SO-spindle coupling is considered important for sleep-dependent memory consolidation. In the current study, using the same N2/3-derived detection thresholds described above, we found that SO-spindle coupling occurred most frequently during N2/3 sleep (2.46 ± 0.06 events/min). Coupling density was significantly lower in N1 (0.75 ± 0.05 events/min, t<sub>(106)</sub> = 23.54, p < 1e-4) and REM sleep (0.43 ± 0.04 events/min, t<sub>(106)</sub> = 31.24, p < 1e-4; Fig. 2b, Table S2-S4), consistent with the expected predominance of SO-spindle coupling in NREM sleep (Ngo et al., 2013; Staresina et al., 2015). As with the individual SO and spindle detections, coupling events detected in N1 and REM were retained only for descriptive stage-wise reporting (see Table S2, S4). They were not used to support physiological claims about SO-spindle coupling in these stages, and they were not entered into the EEG-informed fMRI analyses. All subsequent fMRI GLM and PPI analyses were restricted to N2/3 sleep.”

      “After extracting all N2/3 EEG epochs in which SO-spindle coupling occurred, we analysed their spectral and phase characteristics. The spindles were most likely to occur slightly before the UP-state peak of SOs (Fig. 2a, e), aligning with results from both animal studies (Maingret et al., 2016) and human research (Staresina et al., 2015). In our data, this pattern was consistent across subjects (Fig. 2d, Rayleigh test: z = 9.51, p < 1e-4), with the peak of the spindle aligned at an SO phase of −41.61 ± 0.86° (the SO UP-state peak is 0°).”

      Fig. 2 legend, Page 10, Line 202-205

      “b, SO-spindle coupling density across sleep stages, using SO and spindle detections obtained with fixed N2/3-derived thresholds. Coupling events in N1 and REM are shown only for descriptive comparison. The EEG-informed fMRI analyses used N2/3 coupling events only.”

      Results, Page 11-12, Lines 242-247

      “Second, because the absolute number of detected SO-spindle coupling events depends on the SO detection threshold, we examined whether the main EEG-fMRI results were sensitive to this parameter. To this end, we varied the SO percentile threshold and reconstructed the EEG-informed GLM at each level. Hippocampal activation during SO-spindle coupling remained significant across a range of thresholds (71st - 80th percentile; Fig. S6).”

      Methods, Page 25-26, Lines 567-575

      “Detection of SOs. Data were first bandpass-filtered between 0.16 and 1.25 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). After identifying all positive-to-negative zero crossings, potential SOs were defined based on the interval between consecutive zero crossings, ranging from 0.8 s to 3 s. For each potential SO, we calculated the amplitude range as the peak minus the trough. For each participant, the amplitude threshold was defined as the 75th percentile of candidate SO amplitude ranges observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. Only candidates exceeding this threshold were labelled as SOs, following previous work (Schreiner et al., 2021).”

      Methods, Page 26, Lines 576-583

      “Detection of sleep spindles. Detection of sleep spindles. Data were bandpass-filtered between 12 and 16 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). The root mean square (RMS) of the filtered signal was computed with a 200 ms sliding time window. For each participant, the spindle threshold was defined as the 75th percentile of RMS values observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. RMS segments exceeding this threshold for 0.5 s to 3 s were identified as spindles (Staresina et al., 2015).”

      Methods, Page 26, Lines 584-591

      “Detection of SO-spindle couplings. From the detected SOs and spindles, we identified the peak time of each spindle. Within each SO interval, we checked whether a spindle peak occurred; if so, that SO was labelled as an SO-spindle coupling event. For descriptive stage-wise summaries, coupling events were assigned to the sleep stage of the corresponding SO trough. For every SO-spindle coupling event, an epoch was created time-locked to the SO trough as the central reference, following Schreiner et al. (2021). We extracted data in a [−4 s to 4 s] window around this point, forming the epoch for each coupling event. For the EEG-informed fMRI analyses, only SO, spindle and SO-spindle coupling events detected during N2/3 sleep were used.”

      Methods, Page 26-27, Lines 592-604

      “The detection procedures described above were developed primarily for N2 and N3 sleep, where SOs, spindles and their coupling are physiologically expected and most reliably observed (Hahn et al., 2020; Helfrich et al., 2019; Helfrich et al., 2018; Ngo, Fell, & Staresina, 2020; Schreiner et al., 2022; Schreiner et al., 2021; Staresina et al., 2015; Staresina et al., 2023). Because percentile-based thresholds can otherwise force the detector to label events in every sleep stage, we did not estimate separate thresholds within N1 or REM. Instead, for each participant, all SO and spindle thresholds were defined from N2/3 sleep and then applied uniformly across the recording. Tables S1 and S3 report detailed statistical information on sleep rhythm and N2/3 events detection. The N1 and REM events detection reported in Tables S2 and S4, and illustrated in Fig. S2-S4, should therefore be interpreted as descriptive detector outputs under this fixed N2/3-derived criterion, rather than as evidence for canonical N2/3 SOs, spindles or physiological SO-spindle complexes in those stages. These detections were not used in the EEG-informed fMRI GLM or PPI analyses, which were restricted to N2/3 sleep.”

      Re 10: The rationale for using a lateralized frontal electrode (F3) for both SO (should have been at least bilateral or central) and spindle detection (should have been a centro-parietal electrode) is not convincing. Other EEG-fMRI spindle or SO papers have used a number of frontal (SO) or centro-parietal (spindles) electrodes averaged or even approaches including all EEG electrodes. Searching events with low thresholds at suboptimal recording sites does not dot this highly valuable dataset justice.

      We thank the reviewer for this important comment. We agree that this choice is more sensitive to frontal SOs than to the centro-parietal fast spindle component. Our choice of F3 was driven by the practical constraints of prolonged nocturnal EEG-fMRI recordings. In our MR-compatible EEG setup, FCz was used as the online reference. Central electrodes close to FCz can have reduced signal contrast relative to the reference, and electrodes near the vertex are also more vulnerable to prolonged pressure against the MRI head coil when participants sleep supine for several hours. In this setting, frontal electrodes provided more stable signal quality across the recording. Because the EEG events were used primarily as temporal markers for fMRI modelling, our priority was to obtain reliable event timing during N2/3 sleep rather than to estimate the full scalp topography of SOs and spindles.

      We also agree with your concern that this valuable dataset would ideally be analysed with multichannel detection strategies. To test whether the main result depended on the single lateralised F3 site, we repeated the main EEG-informed fMRI analysis using Fz, a midline frontal electrode. The result was unchanged. Hippocampal activation during SO-spindle coupling remained significant when events were detected from Fz, t<sub>(106)</sub> = 2.47, p = 0.0076, closely matching the original F3-based result, t<sub>(106)</sub> = 2.50, p = 0.0070. This control analysis does not remove the limitation that centro-parietal fast spindles may be underrepresented, and we do not claim that it does. It does show, however, that the main hippocampal fMRI finding is not driven by idiosyncratic detections from one lateralised frontal electrode. We have made this clearer in the revised manuscript.

      Finally, your concern about low thresholds is also important. As described in our response above, the revised analysis now defines SO and spindle thresholds from N2/3 sleep and applies these thresholds uniformly for descriptive comparisons across stages. We also tested the robustness of the hippocampal fMRI result across SO detection thresholds, and the effect remained significant across the 71st to 80th percentile range. We have therefore narrowed the interpretation in the revised manuscript. The main EEG-fMRI result reflects BOLD activity associated with frontal-channel-detected SO-spindle coupling during N2/3 sleep, rather than a full multichannel characterisation of all SO and spindle topographies.

      Results, Page 12, Lines 247-250

      “Third, to test whether the results depended on the use of a single lateralised frontal electrode, we repeated the EEG-informed fMRI GLM using events detected from Fz. Hippocampal activation during SO-spindle coupling again remained significant (t<sub>(106)</sub> = 2.47, p = 0.0076), closely matching the original F3-based result (t<sub>(106)</sub> = 2.50, p = 0.0070).”

      Discussion, Page 18, Lines 380-394

      “Third, sleep oscillation detection was based on a single frontal electrode. This choice improved signal stability and event timing in the prolonged simultaneous EEG-fMRI setting, but it did not exploit the full multichannel EEG information and cannot characterise the full spatial distribution of SOs and spindles. In particular, F3-based detection may be more sensitive to frontal SOs and frontal sigma activity than to the centro-parietal fast spindle component. We therefore interpret the EEG-informed fMRI results as reflecting BOLD activity associated with frontal-channel-detected SOs, spindles, and their coupling during N2/3 sleep. Future studies using multichannel or source-informed detection strategies, with separate treatment of slow and fast spindles, will be better suited to capture the spatial dynamics of these sleep oscillations. Fourth, the use of large anatomical ROIs may mask subregional contributions of specific thalamic nuclei or hippocampal subfields. Finally, without a memory task, we cannot establish a direct behavioral link between sleep-rhythm-locked activation and memory consolidation. Future studies combining ultra-high-field fMRI or iEEG with cognitive tasks, as well as multichannel or source-informed detection strategies that separately characterize slow and fast spindles, will be better suited to refine our understanding of subregional network dynamics and the functional significance of sleep oscillations.”

      Methods, Page 25, Lines 556-566

      “It is worth noting that the primary aim of EEG rhythm detection was to identify reliable event times for EEG-informed fMRI modelling. Detection was performed on the F3 electrode because this channel provided stable signal quality during prolonged nocturnal EEG-fMRI recordings. In our MR-compatible EEG setup, FCz was used as the online reference. Central electrodes close to this reference, and electrodes near the vertex that were in prolonged contact with the head coil during supine sleep, were more susceptible to reduced signal contrast, impedance drift, and pressure-related degradation of electrode-scalp contact. We therefore used F3 as a pragmatic choice to maximize reliable event timing in N2/3 sleep. This choice was not intended to characterise the full scalp topography of SOs or spindles, and it may underrepresent the centro-parietal fast spindle component. As a sensitivity analysis, we repeated the main EEG-informed fMRI GLM using Fz, a midline frontal electrode, with the same detection and modelling procedure.”

      Re 7: It is not clear to me why/how larger voxels would reduce susceptibility-related distortions and partial volume effects. Usually, the opposite is true. This should be elaborated.

      What we meant was that we chose a relatively large voxel size to preserve signal-to-noise ratio and whole-brain coverage within a feasible repetition time for a long overnight EEG-fMRI protocol. This choice is useful for maintaining BOLD sensitivity in sleep recordings, where head motion, physiological noise, and participant comfort are major practical constraints. We agree that it may not be accurate to describe it as reducing susceptibility-related distortion or partial volume effects.

      We have rewritten the Methods to state this trade-off directly. The voxel size of 3.5 × 3.5 × 4.2 mm<sup>3</sup> allowed whole-brain coverage with a TR of 2000 ms, which was important for modelling sleep-rhythm-related BOLD responses across the whole brain during prolonged nocturnal recordings. A smaller voxel size would have improved spatial specificity, but would also have required either a longer TR, reduced brain coverage, or lower SNR, none of which would have been ideal for the present EEG-fMRI sleep design. We now explicitly acknowledge the cost of this choice.

      Methods, Page 21 Lines 453-463

      “For the functional scans, whole-brain images were acquired using a T2*-weighted gradient echo-planar imaging (EPI) sequence sensitive to the BOLD contrast. The sequence parameters were as follows: 33 slices in interleaved ascending order, TR = 2000 ms, TE = 30 ms, voxel size = 3.5 × 3.5 × 4.2 mm<sup>3</sup>, FA = 90°, matrix = 64 × 64, gap = 0.7 mm. A relatively large voxel size was chosen to preserve signal-to-noise ratio while maintaining whole-brain coverage within a feasible repetition time. This compromise was important for the prolonged overnight EEG-fMRI sleep protocol, where head motion, physiological noise, participant comfort, and sustained acquisition stability are substantial practical constraints (Bodurka et al., 2007; Laufs et al., 2008). A smaller voxel size would have improved spatial specificity, but would have required either a longer repetition time, reduced brain coverage, or lower signal-to-noise ratio.”

      Reviewer #2 (Public review):

      In this study, Wang and colleagues aimed to explore brain-wide activation patterns associated with NREM sleep oscillations, including slow oscillations (SOs), spindles, and SO-spindle coupling events. Their findings reveal that SO-spindle events corresponded with increased activation in both the thalamus and hippocampus. Additionally, they observed that SO-spindle coupling was linked to heightened functional connectivity from the hippocampus to the thalamus, and from the thalamus to the medial prefrontal cortex-three key regions involved in memory consolidation and episodic memory processes.

      This study's findings are timely and highly relevant to the field. The authors' extensive data collection, involving 107 participants sleeping in an fMRI while undergoing simultaneous EEG recording, deserves special recognition. If shared, this unique dataset could lead to further valuable insights.

      Thank you for this encouraging assessment. We appreciate your recognition of the effort involved in collecting this simultaneous EEG-fMRI sleep dataset. Below, we respond directly to your remaining concern.

      Comments on revisions:

      The authors' efforts in revising the manuscript and addressing the reviewers' comments are certainly commendable. However, I remain concerned about potential issues in detecting sleep-related oscillations (SOs, spindles, and consequently coupled SO-spindle events), which may arise due to suboptimal parameter selection or inaccurate sleep staging, potentially impacting all subsequent analyses.

      A review of Supplementary Tables 1-4 reveals an unusually high number of detected SOs and spindles during sleep stage N1 and REM sleep. While the authors correctly note that a percentile-based detection approach will always identify a certain number of events across sleep stages, the particularly high counts in N1 and REM are concerning. To mitigate the limitations of this method, the authors could have performed event detection independently of sleep stages (i.e., across the entire dataset for each participant) and subsequently assigned the detected events to the corresponding sleep stages. If the event counts in N1 and REM remained disproportionately high, this would indicate a fundamental issue with the detection procedure.

      In the previous version, thresholds were estimated separately within each sleep stage. As you point out, this can force the detector to identify a relatively large number of SOs and spindles in N1 and REM, even when those waveforms should not be interpreted as canonical N2/3 events.

      We have therefore revised the detection procedure so that event detection is no longer based on separate stage-wise thresholds. Following the logic of your suggestion, we first defined a fixed threshold for each participant and then assigned the detected events to their corresponding sleep stages afterwards. We used N2/3 sleep to define the SO and spindle thresholds because this is the stage in which these events are physiologically expected and most reliably observed. These same N2/3-derived thresholds were then applied unchanged to N1 and REM. This avoids the circularity of forcing a percentile-defined number of detections within each sleep stage.

      With this revised procedure, detections outside N2/3 are clearly lower than those in N2/3. The mean densities are 2.95 SOs/min, 2.71 spindles/min, and 0.75 coupling events/min in N1, and 2.07 SOs/min, 1.81 spindles/min, and 0.43 coupling events/min in REM. We also agree with you that the remaining detections in N1 and REM should not be interpreted as physiological equivalents of canonical N2/3 SOs, spindles, or SO-spindle complexes. We now state this explicitly in the manuscript. They are reported only as descriptive detector outputs under the fixed N2/3-derived criterion. And we have revised all relevant sections of the manuscript, including “[Results, Page 6-7 Lines 134-148]; [Fig. 1e]; [Results, Page 9 Lines 175-191]; [Fig. 2b]; [Methods, Page 25-27, Lines 567-604]; [Fig. S2-S4]; [Table S2, S4].”

      We also would like to clarify our sleep staging procedure. The sleep staging was first performed using an established automated algorithm, YASA toolkit (Vallat & Walker, 2021), and then manually reviewed by two sleep experts. More importantly for the central results, all EEG-informed fMRI GLM and PPI analyses were restricted to N2/3 sleep. Thus, N1 and REM detections did not enter the event regressors used for the main fMRI analyses and do not affect the interpretation of the hippocampal or thalamic findings.

      Finally, we agree that the absolute number of SO-spindle coupling events depends on the detection threshold. We therefore tested whether the main fMRI result depended on the specific SO threshold. Hippocampal activation during SO-spindle coupling remained significant when the SO detection threshold was varied between the 71st and 80th percentiles, as shown in Fig. S6. We have made this clearer in the revised manuscript.

      Results, Page 6-7 Lines 134-148

      “Each sleep stage is characterised by distinct spectral properties and rhythmic waveforms, serving as physiological markers (Fig. 1c). Because SO and spindle detection relies on amplitude-based percentile thresholds, we avoided estimating separate thresholds within each sleep stage. Instead, for each participant, the SO and spindle thresholds were defined from N2/3 sleep only, where these rhythms are most abundant and physiologically expected, and the same fixed thresholds were then applied to N1 and REM for descriptive comparison.”

      “Under this fixed N2/3-derived thresholding, detected SOs and spindles were larger and more frequent in N2/3 than in N1 or REM. SO and spindle amplitudes were significantly higher during N2/3 sleep (SO: 25.59 ± 1.49 μV; spindle: 7.39 ± 0.27 μV) than during N1 (SO: 20.15 ± 2.32 μV; spindle: 5.23 ± 0.27 μV) and REM sleep (SO: 19.84 ± 1.22 μV; spindle: 5.60 ± 0.22 μV; all p < 1e-4; Fig. 1e, Fig. S2). The corresponding event densities showed the same pattern, with 9.64 ± 0.25 SOs/min and 4.19 ± 0.10 spindles/min in N2/3, compared with 2.95 ± 0.16 SOs/min and 2.71 ± 0.14 spindles/min in N1, and 2.07 ± 0.17 SOs/min and 1.81 ± 0.14 spindles/min in REM (all p < 1e-4). We therefore report detections in N1 and REM only as descriptive outputs of the detector under a fixed N2/3-derived criterion, rather than as physiological equivalents of canonical N2/3 SOs or spindles.”

      Fig. 1 legend, Page 8, Line 166-172

      “e, Amplitudes (μV) of detected SOs (left) and spindles (right) across sleep stages. SO and spindle detection thresholds were defined from N2/3 sleep within each participant and then applied unchanged to N1 and REM for descriptive comparison. Detections in N1 and REM should therefore be interpreted as detector outputs under this fixed N2/3-derived criterion. The SO amplitudes were measured from the 0.16-1.25 Hz filtered EEG data, and spindle amplitudes were measured from the 12-16 Hz filtered EEG data. Each dot represents an individual participant. Error bars indicate SEM. *** p < 0.001.”

      Results, Page 9 Lines 175-191

      “SO-spindle coupling is considered important for sleep-dependent memory consolidation. In the current study, using the same N2/3-derived detection thresholds described above, we found that SO-spindle coupling occurred most frequently during N2/3 sleep (2.46 ± 0.06 events/min). Coupling density was significantly lower in N1 (0.75 ± 0.05 events/min, t<sub>(106)</sub> = 23.54, p < 1e-4) and REM sleep (0.43 ± 0.04 events/min, t<sub>(106)</sub> = 31.24, p < 1e-4; Fig. 2b, Table S2-S4), consistent with the expected predominance of SO-spindle coupling in NREM sleep (Ngo et al., 2013; Staresina et al., 2015). As with the individual SO and spindle detections, coupling events detected in N1 and REM were retained only for descriptive stage-wise reporting (see Table S2, S4). They were not used to support physiological claims about SO-spindle coupling in these stages, and they were not entered into the EEG-informed fMRI analyses. All subsequent fMRI GLM and PPI analyses were restricted to N2/3 sleep.”

      “After extracting all N2/3 EEG epochs in which SO-spindle coupling occurred, we analysed their spectral and phase characteristics. The spindles were most likely to occur slightly before the UP-state peak of SOs (Fig. 2a, e), aligning with results from both animal studies (Maingret et al., 2016) and human research (Staresina et al., 2015). In our data, this pattern was consistent across subjects (Fig. 2d, Rayleigh test: z = 9.51, p < 1e-4), with the peak of the spindle aligned at an SO phase of −41.61 ± 0.86° (the SO UP-state peak is 0°).”

      Fig. 2 legend, Page 10, Line 202-205

      “b, SO-spindle coupling density across sleep stages, using SO and spindle detections obtained with fixed N2/3-derived thresholds. Coupling events in N1 and REM are shown only for descriptive comparison. The EEG-informed fMRI analyses used N2/3 coupling events only.”

      Results, Page 11-12, Lines 242-247

      “Second, because the absolute number of detected SO-spindle coupling events depends on the SO detection threshold, we examined whether the main EEG-fMRI results were sensitive to this parameter. To this end, we varied the SO percentile threshold and reconstructed the EEG-informed GLM at each level. Hippocampal activation during SO-spindle coupling remained significant across a range of thresholds (71st - 80th percentile; Fig. S6).”

      Methods, Page 25-26, Lines 567-575

      “Detection of SOs. Data were first bandpass-filtered between 0.16 and 1.25 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). After identifying all positive-to-negative zero crossings, potential SOs were defined based on the interval between consecutive zero crossings, ranging from 0.8 s to 3 s. For each potential SO, we calculated the amplitude range as the peak minus the trough. For each participant, the amplitude threshold was defined as the 75th percentile of candidate SO amplitude ranges observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. Only candidates exceeding this threshold were labelled as SOs, following previous work (Schreiner et al., 2021).”

      Methods, Page 26, Lines 576-583

      “Detection of sleep spindles. Detection of sleep spindles. Data were bandpass-filtered between 12 and 16 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). The root mean square (RMS) of the filtered signal was computed with a 200 ms sliding time window. For each participant, the spindle threshold was defined as the 75th percentile of RMS values observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. RMS segments exceeding this threshold for 0.5 s to 3 s were identified as spindles (Staresina et al., 2015).”

      Methods, Page 26, Lines 584-591

      “Detection of SO-spindle couplings. From the detected SOs and spindles, we identified the peak time of each spindle. Within each SO interval, we checked whether a spindle peak occurred; if so, that SO was labelled as an SO-spindle coupling event. For descriptive stage-wise summaries, coupling events were assigned to the sleep stage of the corresponding SO trough. For every SO-spindle coupling event, an epoch was created time-locked to the SO trough as the central reference, following Schreiner et al. (2021). We extracted data in a [−4 s to 4 s] window around this point, forming the epoch for each coupling event. For the EEG-informed fMRI analyses, only SO, spindle and SO-spindle coupling events detected during N2/3 sleep were used.”

      Methods, Page 26-27, Lines 592-604

      “The detection procedures described above were developed primarily for N2 and N3 sleep, where SOs, spindles and their coupling are physiologically expected and most reliably observed (Hahn et al., 2020; Helfrich et al., 2019; Helfrich et al., 2018; Ngo, Fell, & Staresina, 2020; Schreiner et al., 2022; Schreiner et al., 2021; Staresina et al., 2015; Staresina et al., 2023). Because percentile-based thresholds can otherwise force the detector to label events in every sleep stage, we did not estimate separate thresholds within N1 or REM. Instead, for each participant, all SO and spindle thresholds were defined from N2/3 sleep and then applied uniformly across the recording. Tables S1 and S3 report detailed statistical information on sleep rhythm and N2/3 events detection. The N1 and REM events detection reported in Tables S2 and S4, and illustrated in Fig. S2-S4, should therefore be interpreted as descriptive detector outputs under this fixed N2/3-derived criterion, rather than as evidence for canonical N2/3 SOs, spindles or physiological SO-spindle complexes in those stages. These detections were not used in the EEG-informed fMRI GLM or PPI analyses, which were restricted to N2/3 sleep.”

      Reviewer #3 (Public review):

      Summary:

      Wang et al., examined the brain activity patterns during sleep, especially when locked to those canonical sleep rhythms such as SO, spindle, and their coupling. Analyzing data from a large sample, the authors found significant coupling between spindles and SOs, particularly during the up-state of the SO. Moreover, the authors examined the patterns of whole-brain activity locked to these sleep rhythms. The authors next investigated the functional connectivity analyses, and found enhanced connectivity between the hippocampus and the thalamus and the medial PFC. These results reinforced the theoretical model of sleep-dependent memory consolidation, such that SO-spindle coupling is conducive for systems-level memory reactivation and consolidation.

      Strengths:

      There are obvious strengths in this work, including the large sample size, state-of-the-art neuroimaging and neural oscillation analyses, and the richness of results. The results now inform hemodynamic neural activity that coincided with SO-spindle couplings.

      Weaknesses:

      My earlier comments were about the inability to make inferences on memory given the lack of memory tasks, and the weakness in using the open-ended cognitive state decoding.

      Comments on revisions:

      The current revision has addressed these major concerns. The authors expanded discussions regarding the theoretical implications of the work in a more nuanced manner.

      Thank you for taking the time to re-evaluate the manuscript. We are pleased that the revised Discussion now reads as more nuanced, especially in relation to the limits of the memory-related interpretation. Your earlier comments helped us sharpen both the claims and the framing, and we are grateful for that.

      References:

      Bergmann, T. O., Mölle, M., Diedrichs, J., Born, J., & Siebner, H. R. (2012). Sleep spindle-related reactivation of category-specific cortical regions after learning face-scene associations. Neuroimage, 59(3), 2733-2742.

      Bodurka, J., Ye, F., Petridou, N., Murphy, K., & Bandettini, P. A. (2007). Mapping the MRI voxel volume in which thermal noise matches physiological noise—implications for fMRI. Neuroimage, 34(2), 542-549.

      Caporro, M., Haneef, Z., Yeh, H. J., Lenartowicz, A., Buttinelli, C., Parvizi, J., & Stern, J. M. (2012). Functional MRI of sleep spindles and K-complexes. Clinical neurophysiology, 123(2), 303-309.

      Czisch, M., Wehrle, R., Stiegler, A., Peters, H., Andrade, K., Holsboer, F., & Sämann, P. G. (2009). Acoustic oddball during NREM sleep: a combined EEG/fMRI study. PloS one, 4(8), e6749.

      Fogel, S., Albouy, G., King, B. R., Lungu, O., Vien, C., Bore, A., Pinsard, B., Benali, H., Carrier, J., & Doyon, J. (2017). Reactivation or transformation? Motor memory consolidation associated with cerebral activation time-locked to sleep spindles. PloS one, 12(4), e0174755.

      Hahn, M. A., Heib, D., Schabus, M., Hoedlmoser, K., & Helfrich, R. F. (2020). Slow oscillation-spindle coupling predicts enhanced memory formation from childhood to adolescence. Elife, 9, e53730.

      Hale, J. R., White, T. P., Mayhew, S. D., Wilson, R. S., Rollings, D. T., Khalsa, S., Arvanitis, T. N., & Bagshaw, A. P. (2016). Altered thalamocortical and intra-thalamic functional connectivity during light sleep compared with wake. Neuroimage, 125, 657-667.

      Helfrich, R. F., Lendner, J. D., Mander, B. A., Guillen, H., Paff, M., Mnatsakanyan, L., Vadera, S., Walker, M. P., Lin, J. J., & Knight, R. T. (2019). Bidirectional prefrontal-hippocampal dynamics organize information transfer during sleep in humans. Nature Communications, 10(1), 3572.

      Helfrich, R. F., Mander, B. A., Jagust, W. J., Knight, R. T., & Walker, M. P. (2018). Old brains come uncoupled in sleep: slow wave-spindle synchrony, brain atrophy, and forgetting. Neuron, 97(1), 221-230. e224.

      Huang, Q., Xiao, Z., Yu, Q., Luo, Y., Xu, J., Qu, Y., Dolan, R., Behrens, T., & Liu, Y. (2024). Replay-triggered brain-wide activation in humans. Nature Communications, 15(1), 7185.

      Ilhan-Bayrakcı, M., Cabral-Calderin, Y., Bergmann, T. O., Tüscher, O., & Stroh, A. (2022). Individual slow wave events give rise to macroscopic fMRI signatures and drive the strength of the BOLD signal in human resting-state EEG-fMRI recordings. Cerebral Cortex, 32(21), 4782-4796.

      Laufs, H., Daunizeau, J., Carmichael, D. W., & Kleinschmidt, A. (2008). Recent advances in recording electrophysiological data simultaneously with magnetic resonance imaging. Neuroimage, 40(2), 515-528.

      Maingret, N., Girardeau, G., Todorova, R., Goutierre, M., & Zugaro, M. (2016). Hippocampo-cortical coupling mediates memory consolidation during sleep. Nature Neuroscience, 19(7), 959-964.

      Moehlman, T. M., de Zwart, J. A., Chappel-Farley, M. G., Liu, X., McClain, I. B., Chang, C., Mandelkow, H., Özbay, P. S., Johnson, N. L., & Bieber, R. E. (2019). All-night functional magnetic resonance imaging sleep studies. Journal of neuroscience methods, 316, 83-98.

      Ngo, H.-V., Fell, J., & Staresina, B. (2020). Sleep spindles mediate hippocampal-neocortical coupling during long-duration ripples. Elife, 9, e57011.

      Ngo, H. V., Martinetz, T., Born, J., & Molle, M. (2013). Auditory closed-loop stimulation of the sleep slow oscillation enhances memory. Neuron, 78(3), 545-553.

      Picchioni, D., Horovitz, S. G., Fukunaga, M., Carr, W. S., Meltzer, J. A., Balkin, T. J., Duyn, J. H., & Braun, A. R. (2011). Infraslow EEG oscillations organize large-scale cortical–subcortical interactions during sleep: a combined EEG/fMRI study. Brain research, 1374, 63-72.

      Schabus, M., Dang-Vu, T. T., Albouy, G., Balteau, E., Boly, M., Carrier, J., Darsaud, A., Degueldre, C., Desseilles, M., & Gais, S. (2007). Hemodynamic cerebral correlates of sleep spindles during human non-rapid eye movement sleep. Proceedings of the National Academy of Sciences, 104(32), 13164-13169.

      Schreiner, T., Kaufmann, E., Noachtar, S., Mehrkens, J.-H., & Staudigl, T. (2022). The human thalamus orchestrates neocortical oscillations during NREM sleep. Nature Communications, 13(1), 5231.

      Schreiner, T., Petzka, M., Staudigl, T., & Staresina, B. P. (2021). Endogenous memory reactivation during sleep in humans is clocked by slow oscillation-spindle complexes. Nature Communications, 12(1), 3112.

      Staresina, B. P., Bergmann, T. O., Bonnefond, M., van der Meij, R., Jensen, O., Deuker, L., Elger, C. E., Axmacher, N., & Fell, J. (2015). Hierarchical nesting of slow oscillations, spindles and ripples in the human hippocampus during sleep. Nature Neuroscience, 18(11), 1679-1686.

      Staresina, B. P., Niediek, J., Borger, V., Surges, R., & Mormann, F. (2023). How coupled slow oscillations, spindles and ripples coordinate neuronal processing and communication during human sleep. Nature Neuroscience, 1-9.

      Vallat, R., & Walker, M. P. (2021). An open-source, high-performance tool for automated sleep staging. Elife, 10.

    1. eLife Assessment

      By mining the Logan assemblage of the Sequence Read Archive to reveal substantial papillomavirus diversity, this valuable study establishes a framework for integrating viral discovery with host, geographic, and ecological metadata. The evidence for identifying novel papillomavirus sequences is convincing, supported by large-scale sequence searches, established L1-based typing criteria, phylogenetic analyses, protein annotation, and structural comparisons. The broader host-association and ecological conclusions are more tenuous due to heterogeneous public metadata and uneven sampling, and would benefit from clearer discussion of potential biases, contamination, endogenous viral elements, and alternative explanations. The work should be of interest to a broad range of colleagues in the areas of microbiome research, virology, and viral evolution.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript titled "Petabase-scale papillomavirus discovery" capitalizes on Logan assemblages to identify novel and known PVs.

      This is a brilliant use of Logan assemblages, and this paper highlights the use of this, plus also shows that one person's trash is another one's gold. Superb paper and I applaud the authors for starting with the PVs as easier to identify due to their set of genes coupled with the conserved L1 protein and associated typing for PVs (10% pairwise identity threshold for identification of new PV types).

      Strengths:

      This study highlights the hidden gems in public resources, especially if mined properly. Thanks to Logan assemblages, this is possible, and this manuscript highlights this with their data mining of papillomaviruses, identifying known and novel PVs in pangolins, lizards, fish, and white rhinos.

      Weaknesses:

      None identified

    3. Reviewer #2 (Public review):

      This study applies the Logan assemblage of the Sequence Read Archive to investigate papillomavirus diversity at a large scale. The authors combine sequence similarity searches, phylogenetic analyses, protein annotation, structural comparisons, and metadata integration to identify and characterize novel papillomavirus sequences. Beyond papillomavirus sequence discovery, the study develops a framework for associating viral sequences with host, geographic, and ecological metadata through the integration of both established and recently developed large-scale public sequence repositories. This framework may provide a foundation for similar investigations across other viral taxa.

      An important aspect of the work is the systematic processing and integration of metadata across a large number of sequencing libraries. The approaches developed to aggregate, curate, and interpret host and environmental information may be broadly applicable to future large-scale studies of other viral families. At the same time, interpretations based on metadata-derived ecological and host-association patterns should be considered in the context of the inherent limitations of public sequencing repositories, including uneven taxonomic representation, heterogeneous sampling strategies, laboratory-derived samples, and variable sequencing protocols. The comments below primarily address methodological details, interpretation of ecological patterns, and opportunities to further strengthen the robustness of the analyses and conclusions.

      Pages 136-137: The manuscript focuses primarily on L1-containing contigs. Could the authors provide a summary of libraries containing other PV hallmark genes (e.g., E1 or E2) but lacking full-length L1 sequences? This would help assess how much additional PV diversity may remain inaccessible under an L1-centered framework. Furthermore, the manuscript does not provide a detailed discussion of technical or biological explanations for libraries with detected PV sequences but lacking L1 sequences. Could such cases represent incomplete assemblies, low-abundance infections, highly divergent PVs, or endogenous papillomavirus-derived elements? Clarifying these possibilities would help readers interpret the biological significance of these detections.

      Page 141: The rationale for performing the search in two sequential Logan releases is unclear. Does Logan v1.1 fully supersede v1.0? If so, why was the initial search performed on v1.0 rather than directly on v1.1? Providing a clearer description of the differences between the two database versions and explaining how these differences motivated the two-round search strategy would improve reproducibility.

      The data could have been explored at the libraries' read level. The analysis is currently limited to presence/absence and diversity patterns derived from assembled PV contigs. However, the identified PV-positive libraries provide an opportunity to explore abundance-related metrics. Read counts, coverage estimates, or other measures of sequence representation could be used to characterize PV abundance within libraries, providing additional ecological context and helping distinguish low-level incidental detections from strongly represented infections. In addition, the curated PV sequence dataset generated in this study could serve as a reference for targeted read-mapping analyses. Aligning reads from a subset of libraries classified as PV-negative may help determine whether PV sequences are present below assembly detection thresholds. Such an analysis could provide valuable insights into the sensitivity of assembly-based virus discovery approaches and help establish practical coverage or read-count thresholds for detecting low-abundance papillomaviruses. These results could have important implications for future surveillance, clinical, and environmental studies aimed at PV detection.

      Pages 155-156: Clustering was performed using sequence identity, while host, geographic, and ecological annotations were assigned from representative centroids. How frequently did clusters contain sequences associated with conflicting metadata (e.g., distinct hosts or geographic regions), and how were such cases handled?

      Page 160: The authors state that the 70% query coverage threshold was selected based on an observed bimodal distribution. Could this analysis be shown explicitly (e.g., in a supplementary figure), and could the authors discuss how sensitive the number of novel PV calls could have been to alternative coverage thresholds?

      Pages 167-168: The rationale for clustering highly divergent sequences at 60% nucleotide identity should be explained. Is this threshold associated with established genus-level classifications in this viral family, or was it chosen empirically?

      Pages 167-168: For the 45 sequences lacking nucleotide-level matches, did the authors investigate amino-acid similarity to known PV L1 proteins? Such analyses would help determine whether these sequences represent deeply divergent PVs or potentially more distant viral lineages.

      Pages 181-183: The biological interpretation of host-associated PV diversity may depend on library type. Could the authors summarize the proportion of samples originating from field collections, laboratory animals, cell culture systems, or experimental infections?

      Pages 243-248: Could the observed geographic and ecological patterns be influenced by laboratory-derived samples? Distinguishing field-collected samples from laboratory, captive, or experimental material would strengthen the ecological interpretations.

      Pages 254-261: The biome analyses focus on PV occurrence. An analysis of host composition across biomes would be highly informative and could help disentangle whether observed patterns reflect PV ecology or underlying host distributions.

      Page 294: Statements regarding structural similarity appear to rely primarily on visual comparisons. Could the authors provide quantitative structural alignment metrics (e.g., RMSD, TM-score, DALI score, Foldseek score) to support these conclusions?

      Page 321-324: Given the scale and curation of the dataset, the final case-study section is limited. Broader comparative analyses of gene content, ORF architecture, and composition related to host associations, phylogenetic relationships, and/or ecological variables could provide additional evolutionary insights beyond a small number of illustrative examples.

      Page 333: Given the emphasis on the feasibility of petabase-scale sequence mining, the manuscript would benefit from a more detailed description of the computational resources required. The reported ~10-hour runtime is difficult to interpret without information regarding hardware specifications, CPU-hours, memory requirements, storage footprint, and cloud infrastructure (if used). Such information is important for evaluating the reproducibility and practical applicability of the approach.

      Pages 408-409: Were metadata available regarding viral enrichment procedures, particle purification, or size-selection protocols? Such information could influence the interpretation of PV detection frequencies across library types.

      Pages 415-416: The differentiation between viral and endogenous viral sequences is one of the biggest challenges in viral metagenomics and large-scale data mining for viral sequences. This issue is particularly relevant because the distinction between exogenous and endogenous viral sequences may directly affect estimates of novel PV diversity and inferred host associations. The manuscript acknowledges that papillomavirus sequences recovered from DNA-based libraries may derive from integrated viral DNA. However, there is no systematic analysis addressing the potential contribution of endogenous papillomavirus elements (EVEs) to the reported diversity estimates. Given the large number of host genome sequencing projects represented in the SRA, some detected PV-like sequences may correspond to integrated or fossil viral sequences rather than exogenous viruses. The authors could discuss this possibility more explicitly and provide analyses evaluating the prevalence of integration signatures, disrupted ORFs, host-genome flanking regions, or other indicators that would help distinguish endogenous viral elements from actively circulating PVs.

      Pages 532-533: The study is described as "petabase-scale"; however, the analyses were performed on a pre-assembled and compressed representation of the SRA rather than directly on petabase-scale raw sequencing data. The authors may wish to clarify this distinction and explicitly acknowledge that the computational burden is substantially reduced by the Logan framework and, from this perspective of computational power applied, this study is not in the same context as Serratus and Logan.

      Perspective comment: One of the strengths of this study is the generation of a highly curated papillomavirus protein dataset spanning a broad range of known and newly identified PV diversity. Given the increasing importance of structure-based homology detection in virology, the authors may wish to discuss the potential of this resource for future structure-guided discovery efforts. Recent studies have shown that protein structure prediction and comparison can reveal extremely distant evolutionary relationships that are undetectable at the sequence level. The curated PV dataset generated here could serve as a valuable reference for searching unannotated proteins from metagenomic "dark matter" datasets for structural homologs or convergent folds related to papillomavirus proteins. Such approaches may help identify highly divergent PV lineages or previously unrecognized viral proteins that retain structural similarity despite extensive sequence divergence.

  2. Jul 2026
    1. eLife Assessment

      This valuable study reports that ALDH-abundant cells exhibit stem cell properties and may play a key role in endometrial epithelial development in mice. The data support the main conclusion and are convincing. This work will be of great interest to reproductive biologists and biomedical researchers working on women's reproductive health.

    2. Reviewer #1 (Public review):

      The manuscript by Tang et al. characterizes the expression dynamics and functional roles of aldehyde dehydrogenase 1 activity in uterine physiology. Using a combination of in vivo lineage tracing and cell ablation coupled with organoid culture, the authors propose that Aldh1a1 lineage-marked cells contribute to uterine gland development and epithelial regeneration. The descriptive data will be of interest to reproductive biologists and clinicians and builds on established hypotheses in the field. The manuscript is well written and scientifically sound and the important experimental limitations and interpretation caveats are presented throughout.

      The majority of my initial comments have been adequately addressed within the text.

      Remaining limitations include:

      (1) The impact of tamoxifen injection directly on Aldh1a1 expression in the developing uterus.

      (2) It would be beneficial to demonstrate the degree of cell death following diphtheria toxin treatment 24-48 hours after injection in Tam-treated mice at PND 10. It is not clear as to why the 4-day timepoint was selected. Cells expressing the DTR should begin undergoing apoptosis within several hours after treatment.

    3. Reviewer #2 (Public review):

      Tang et al. investigated the contribution of Aldh1a1+ cells, as putative stem/progenitor cells, to endometrial development, maintenance during the estrous cycle, and postpartum repair in mouse models. They employed in vitro organoid formation and in vivo lineage tracing models coupled with RNA-seq to test the stem-ness of Aldh1a1+ cells. They found that mouse endometrial cells with high ALDH activity (using the ALDEFLUOR assay) formed more and larger organoids and were enriched for stem/progenitor cell gene signatures. Similar results were shown using endometrial cells from a human patient sample. Epithelial ALDH1A1 expression was shown to be hormonally regulated, becoming more restricted to the glands, a putative epithelial stem cell niche, under estrogen stimulation. Using lineage-tracing initiated postnatally/prepubertally, Aldh1a1+ epithelial cells were shown to expand, contributing to both the luminal and glandular epithelium into adulthood, whereas adult initiation of labeling showed expansion of stromal Aldh1a1+ cells but not epithelial. Postnatal ablation of single-labeled Aldh1a1+ epithelial cells resulted in impaired gland development. Lastly, Aldh1a1-lineage traced cells (adult labeled) were present during postpartum endometrial repair as were epithelial/mesenchymal transitional cells.

      Comments on revised version.

      In the revised manuscript, comments have largely been addressed and the manuscript is improved. The authors have tempered their inference of the contribution of ALDH1A1+ cells to endometrial regeneration, but the conclusions are still somewhat overstated. However, the overall work provides valuable insight and strengthens the growing body of literature characterizing endometrial stem/progenitor cells and their function.

    4. Reviewer #3 (Public review):

      Summary:

      Tan et al demonstrated the importance of ALDH-high cells in the epithelial development in the mouse endometrium, and these cells displayed properties of stem cells.

      Strengths:

      The findings are solid, supported and validated through a combination of technical methods. I appreciated this combined use of mouse and human endometrial cells to strengthen the findings. Genomic results from a single-cell sequencing dataset were informative as they depicted the different stages of the estrus cycle during the regeneration process. Verification with immunostainings with various markers made it convincing for readers to visualize the cell's location, progression, and status at different timepoints. Utilizing human endometrial cells further demonstrated that the phenomenon observed in mice can be translated to humans.

      This work will greatly advance the understanding of endometrial regeneration for reproductive biologists.

      Comments on revised version.

      The authors have answered the questions in the revised manuscript, no further comments.

    5. Author response:

      The following is the authors’ response to the original reviews

      eLife Assessment

      This valuable study reports that the ALDH-abundant cells display stem cell properties and may play a key role in the endometrial epithelial development in the mouse. The data supporting the main conclusion are solid, although further improvements are needed to strengthen the conclusions. This work will be of great interest to reproductive biologists and biomedical researchers working on women's reproductive health.

      We thank the reviewers and editor for their critical reading and assessment of our manuscript. We carefully considered each of the points raised by the reviewers. In this document and in the edited manuscript and figures, we have carefully addressed each of the comments and requested modifications. In light of these changes, we expect that you will find that the manuscript has improved.

      We indicate our responses to the reviewers below in blue font and highlight the changes in the manuscript using the line numbers corresponding to the tracked version of the revised document.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript by Tang et al. characterizes the expression dynamics and functional roles of aldehyde dehydrogenase 1 activity in uterine physiology. Using a combination of in vivo lineage tracing and cell ablation coupled with organoid culture, the authors propose that Aldh1a1 lineage-marked cells contribute to uterine gland development and epithelial regeneration. The descriptive data will be of interest to reproductive biologists and clinicians and will build on established hypotheses in the field. The manuscript is well written and scientifically sound; however, several experimental limitations and interpretation caveats should be addressed.

      We thank the reviewer for their comments and expert assessment of our paper.

      (1) The methods surrounding the passage number and duration of culture following sorting prior to transcriptomic profiling should be clarified in the figure legends. Related to this, the representative images in Figures 1D and 1E do not appear consistent with the quantification presented in Figures 1F-H and should be reconciled.

      Thanks for this comment. We have now clarified this in the Figure 1 legend as follows,

      LINES 1026-1029: “Organoid formation assay performed immediately after luminal epithelial cell isolation and by plating equal numbers of viable ALDH<sup>LO</sup> (D) and ALDH<sup>HI</sup> (E) epithelial cells. ALDH<sup>LO</sup> and ALDH<sup>HI</sup> organoids were cultured for two weeks and passaged once prior to the organoid formation assays and transcriptomic analyses.”

      Regarding the second comment, we recognize that the images we showed may not have been the most representative of our quantification. As such, we replaced them with the organoid images so that they better reflect the quantification outlined in Figure 1F-H.

      (2) The conclusion that ALDH1A1+ cells are enriched in populations with stem cell characteristics relies primarily on transcriptomic analysis. Protein-level co-localization should be performed to strengthen this claim.

      We thank the reviewer for this comment. Unfortunately, the antibodies for many of these stem cell markers (such as LGR5, AXIN2, and SUSD2) are not well-suited for immunostaining. Others that have been proposed in human and are amenable to immunostaining are not suitable markers for mouse endometrial stem cells (such as CDH2). We hope that by showing that ALDH1A1 is expressed in patterns that are similar to the previously published stem cell markers LGR5 and AXIN2 (i.e., throughout the epithelium in the developing uterus and subsequently enriched in the tips of the endometrial glands of adult mice), along with transcriptomic studies, we can demonstrate its utility as a marker for mouse endometrial stem cells.

      (3) The overlap of 19 genes between the data set here and AXIN2 HI data is presented as evidence of shared stemness identity, but no statistical assessment of this overlap is provided. A hypergeometric test should be performed to determine whether this overlap is greater than expected by chance.

      Thank you for this suggestion. We have performed a hypergeometric test and determined that the reported shared genes between the two datasets are greater than is expected by chance. We have updated the results section to state the following:

      Lines 137-140: "We determined that the overlap between ALDH<sup>HI</sup> and Axin2<sup>+</sup> stemness marker genes was significantly greater than expected by chance for both upregulated (21/346 genes, 1.81-fold enrichment, p = 0.0067) and downregulated (19/674 genes, 1.67-fold enrichment, p = 0.021) gene sets (hypergeometric test, universe = 23,182 genes)."

      (4) The impact of tamoxifen injection on Aldh1a1 expression should be characterized in the neonatal uterus, as tamoxifen itself has known estrogenic activity that could confound interpretation of the lineage tracing results at early postnatal timepoints.

      Although we took measures to control for this possibility by using multiple time-points and models to trace the impact of Aldh1a1<sup>+</sup> cells in development and adulthood, we recognize the importance of this comment and acknowledge that this is a limitation in the design of our study. We have included the following text to the Discussion acknowledging this point:

      Lines 433-441: “Given the well-documented impacts of tamoxifen for lineage tracing studies, it is imperative to use doses of tamoxifen that will minimize estrogenic impacts and result in off-target effects (Rios et al., 2016). This often requires administration at doses that will achieve maximal recombination of the desired gene, while ensuring that the potential deleterious impacts of tamoxifen are minimized (Chen et al., 2023; Pimeisl et al., 2013). The cre/ERT2 tamoxifen inducible model is widely used to study uterine biology where it serves as a useful tool to interrogate the spatiotemporal impact of key genes, either through inactivation or for lineage tracing. Despite its widely documented utility across many tissue types and developmental timepoints, the use of tamoxifen and its impacts on the endometrium remain a limitation of our study, which we tried to address by implementing multiple timepoints, doses, and orthogonal assays in our experimental design.”

      (4b) Related to this, while low-dose tamoxifen is shown to label individual cells within 24 hours of injection, the translation dynamics of the label following Cre-mediated recombination can require up to 72 hours. The presence of only a few labeled clones at PND8 but multiple separate clones per cross-section at later timepoints warrants discussion and may reflect labeling kinetics rather than clonal expansion.

      The reviewer raises an important point. We agree that the 72hr-translation kinetics of the cre-mediated recombination is a legitimate consideration for interpreting our data and we have added the text below to the Discussion section acknowledging this point.

      We have addressed this by adding the following text to the discussion:

      Lines 417-422: We hypothesized that the singly labeled cells observed from one day tracing experiments expanded in a clonal fashion during the various timepoints we measured. We note that the translation kinetics of the labeled cells following cre-mediated recombination may contribute to the limited labeling observed at PND8/PND15 and there is a potential for delayed labeling of cells between 24 and 72 hours of tamoxifen administration. However, the continuous increase in labeled cells at the subsequent timepoints favors our interpretation of clonal expansion as the primary explanation.

      (5) It would strengthen the in vivo ablation data to validate the degree of cell death following diphtheria toxin treatment directly. It is possible that a general decrease in cell number rather than specific loss of a stem cell population is responsible for the observed reduction in gland number and FOXA2 expression (Tongtong et al 2017).

      We agree that this is an important control to incorporate into our experimental design. To rule out this possibility, we performed immunohistochemistry of cleaved caspase 3 in the uterine tissues of DTR<sup>flox/flox</sup> and DTR<sup>flox/flox</sup>;Aldh1a1<sup>cre/ERT2</sup> mice 4 days after administration of diphtheria toxin. The results indicate similar levels of cleaved caspase 3 detection in both genotypes, suggesting that the decrease in FOXA2+ cells is not due to non-specific cell death, but rather the result of ALDH1A1<sup>+</sup> cells. These data and the following text have been added to the manuscript:

      Lines 320-324: “We determined that the decreased in FOXA2<sup>+</sup> cells in the experimental mice was not the result of non-specific DT-mediated cell death, as similar levels of cleaved caspase 3-positive cells were detected in the DT-treated control ROSA26<sup>DTR/DTR</sup> and ROSA26<sup>DTR/DTR</sup>;Aldh1a1<sup>cre/ERT2/+</sup> mice 4 days post-diphtheria toxin administration (Figure S3G-H’).”

      (6) The lineage tracing data in the postpartum endometrium demonstrate that Aldh1a1-marked cells are present during regeneration, but it remains unclear whether these cells are preferentially activated or expanded in response to tissue injury. Coupling these studies with diphtheria toxin-mediated ablation during active regeneration would more directly test the proposed regenerative role of this population.

      This is a great point and one that we would be very interested in pursuing as follow-up studies in our future work. Regretfully, due to the long generation time and experimental procedures associated with these proposed studies, we are not able to include these experiments in the current manuscript. Thus, we have changed our wording and conclusions throughout the manuscript to be less definitive in terms of the role of Aldh1a1 in regeneration, since this will be the focus of future studies.

      The contribution of stromal Aldh1a1 lineage-positive cells is underexplored in the discussion, given the lineage tracing data showing stromal labeling across multiple timepoints and its potential relevance to mesenchymal-to-epithelial transition.

      Thank you for the suggestion. We have now expanded this section in the Discussion to include the following:

      Lines 496-504: We also found ALDH1A1<sup>+</sup> stromal cells were more prevalent when tracing began in adult mice. Other studies have shown that mesenchymal cells contribute to endometrial regeneration in the postpartum phase or after induced menses through a process of MET (Cousins et al., 2014; Kirkwood et al., 2022; Li et al., 2025). Similarly, lineage tracing studies have shown that MET is an active process and contributes to epithelial cell regeneration in the post-partum phase (Huang et al., 2012; Patterson et al., 2013). Although this is an area of active investigation in the field, with some contradicting reports, it is plausible to hypothesize that endometrial tissue has the capacity to undergo wound-healing and regeneration via several mechanisms (Ang et al., 2023; Ghosh et al., 2020). The process of MET in wound healing is widely documented in other organs, such as the kidney, liver and lung, where MET is associated with depletion of the resident epithelial cell pool (Bi et al., 2012; Niayesh-Mehr et al., 2024; Zeisberg et al., 2005).

      Finally, the word 'control' may overstate the functional evidence presented. 'Contribute' may be more accurate given the partial and context-dependent nature of the phenotypes observed.

      We agree with the reviewer’s point that control may overstate the evidence that we provide in the manuscript. To reflect this, we have edited the manuscript title and text to address this suggestion.

      Reviewer #2 (Public review):

      Tang et al. investigated the contribution of Aldh1a1+ cells, as putative stem/progenitor cells, to endometrial development, maintenance during the estrous cycle, and postpartum repair in mouse models. They employed in vitro organoid formation and in vivo lineage tracing models coupled with RNA-seq to test the stem-ness of Aldh1a1+ cells. They found that mouse endometrial cells with high ALDH activity (using the ALDEFLUOR assay) formed more and larger organoids and were enriched for stem/progenitor cell gene signatures. Similar results were shown using endometrial cells from a human patient sample. Epithelial ALDH1A1 expression was shown to be hormonally regulated, becoming more restricted to the glands, a putative epithelial stem cell niche, under estrogen stimulation. Using lineage-tracing initiated postnatally/prepubertally, Aldh1a1+ epithelial cells were shown to expand, contributing to both the luminal and glandular epithelium into adulthood, whereas adult initiation of labeling showed expansion of stromal Aldh1a1+ cells but not epithelial. Postnatal ablation of single-labeled Aldh1a1+ epithelial cells resulted in impaired gland development. Lastly, Aldh1a1-lineage traced cells (adult labeled) were present during postpartum endometrial repair as were epithelial/mesenchymal transitional cells.

      This study addresses an important area of research in the field of endometrial stem/progenitor cell biology. The authors are commended for their use of multiple complementary methods, including lineage tracing, DTR-mediated cell ablation, organoid assays, and RNA-seq in mouse and human models to assess the stem-like nature of Aldh1a1+ cells. The data support the stem/progenitor phenotype of Aldh1a1+ epithelial cells during endometrial development; however, there are noted discrepancies between organoid formation assays and lineage tracing experiments regarding the stemness of Aldh1a1+ epithelial cells in adults. Specifically, organoids were generated from adult cells and demonstrated in vitro stem cell activity; however, in vivo lineage-tracing of adult cells either during the estrous cycle or postpartum repair does not show expansion of Aldh1a1+ cells, suggesting they do not have stem/progenitor activity. Additionally, the stem-ness of epithelial vs stromal Aldh1a1+ cells is confounded in the study because epithelial cells were not purified for organoid experiments, epithelial cells were not exclusively lineage-traced as stromal cells were also labeled, and mesenchymal-epithelial transition was suggested to occur during postpartum repair. The following specific comments are presented to detail these concerns:

      We thank the reviewer for their critical reading of our manuscript and constructive comments.

      (1) The statement in the brief summary, "...critical for lifelong endometrial regeneration," is not supported by the data provided.

      We have edited the brief summary to exclude this statement, it now reads as follows:

      Lines 4-5: “We uncover ALDH1A1<sup>+</sup> cells as a group of hormone sensitive stem cells contributing to endometrial development and regeneration.”

      (2) AlDH1A1 is not restricted to the endometrial epithelium, and epithelial cells were not purified by flow cytometry for experiments in Figure 1. Figure 2 clearly shows the presence of mesenchymal cells, even using the described method for enriching for epithelial cells. Therefore, contaminating mesenchymal cells with high ALDH activity may confound the experimental results in Figure 1, either through promoting epithelial cell growth or through MET. The authors should provide clear evidence of epithelial purity in organoid experiments or that mesenchymal cells are not contained in the ALDHhi population. These comments also apply to the human organoid experiments in Figure 7.

      We thank the reviewer for raising this important point. Our group has been using the enzymatic method to routinely separate epithelial from stromal cell populations from the mouse uterus (see references dating back to 2015, PMID 26721398, 28324064, 34099644). In these experiments we typically obtain >98% purity in the epithelial and stromal cell compartments, respectively. We can directly observe this purity in the immunofluorescence images shown I Author response image 1 and Author response image 2, where mouse endometrial epithelial cells and stromal cells were enzymatically separated and immunostained with E-cadherin and vimentin antibodies to detect epithelial and mesenchymal cells in both cell preparations. The images show very few contaminating epithelial and stromal cells in either cell preparation. We have observed similar results when preparing epithelial and stromal cell preparation from the human endometrium, where the epithelial cell organoids display high purity with ~100% epithelial cell expression when we perform immunostaining.

      Author response image 1.

      Purity of mouse endometrial epithelial cells obtained via enzymatic and mechanical dissociation. A-B) Shows the epithelial (A) and stromal (B) cells plated on glass coverslips and immunostained with an epithelial cell marker (cytokeratin 8, red), a stromal cell marker (vimentin, green), and DAPI.

      Author response image 2.

      Human endometrial epithelial organoids were fixed and immunostained with cytokeratin 8 (green) and DAPI. The images are typical for our epithelial cell cultures and demonstrate that all epithelial cells are CK8-positive.

      (3) Lines 186-187: Susd2 was increased in EpSC clusters, yet this is a mesenchymal stem/progenitor marker in humans. The authors should discuss the implications of this.

      We thank the reviewer for highlighting this. We have now included the following in our Discussion to address this point:

      Lines 527-532: Clustering with this population of EpSCs were Susd2<sup>+</sup> cells, which are well-characterized mesenchymal progenitors that are enriched in the perivascular regions of the human endometrium (Darzi et al., 2016; Khanmohammadi et al., 2021). The presence of Susd2<sup>+</sup> cells, while unexpected in an epithelial stem cell niche, could indicate the presence of a transitional mesenchymal or perivascular cell that is differentiating into epithelium. Evidence for both mesenchymal and Nestin2<sup>+</sup> pericytes have been recently described in the mouse endometrial epithelium (Kirkwood et al., 2022; Li et al., 2025).

      (4) In Figure 5, RFP+ epithelial cells should be quantified as in previous figures to substantiate the statement in lines 279-280, "At PPD5, the proportion of RFP+ epithelial cells had expanded relative to PPD1 and PPD3 (Figure 5E-E')." Especially because in the low mag images (C-E), RFP+ epithelial cells appear to be most abundant at PPD1 and decrease at PPD3 and PPD5, suggesting that they may not be involved in endometrial regeneration/repair (contradicting the interpretation in line 285). Further, if there is in fact a decrease over postpartum repair, then regeneration should be removed from the title of the manuscript. RFP+ stromal cells should also be quantified.

      We appreciate this reviewer’s comment and agree that as stated, the conclusion is not fully supported by the data. To address this comment, we have edited the results so that they clearly indicate the results and remove any ambiguity:

      As requested, we quantified the number of RFP+ stromal and epithelial cells during the postpartum phase and noted that RFP+ cells were prominent in the stromal compartment of the endometrium. While RFP+ epithelial were also observed during these timepoints, they were less abundant than RFP+ stromal cells. Because the number of RFP+ cells did not significantly change over the postpartum phases in neither the stromal nor epithelial compartment, we have modified our conclusion to state that ALDH1A1+ cells are transiently detected in the regenerating endometrium.

      Results:

      Lines 287-294: “By analyzing the uterine tissues near the placental detachment site, we observed that RFP positive cells were prominent in the endometrial stromal cells that were adjacent to the luminal epithelium (Figure 5C-C’, green arrows). RFP<sup>+</sup> cells were also observed in the stromal cells near the placental detachment sites at PPD1 and PPD3 (Figure 5D’-E’, red & blue arrows) and in limited luminal epithelial cells (Figure 5D”,E”). Quantification of RFP+ cells throughout these postpartum phases indicated that stromal cells had more frequent ALDH1A1<sup>+</sup> stromal cells (360 ± 103, PPD1, n=3; 217 ± 107, PPD3, n=3; 254 ± 32, PPD5, n=4) than ALDH1A1<sup>+</sup> epithelial cells in the regenerating endometrium (65 ± 65, PPD1, n=3; 20 ± 10, PPD3, n=3; 114.25 ± 39, PPD5, n=4) (Figure S4).”

      Discussion:

      Lines 512-520: “We also noted that a majority of ALDH1A1<sup>+</sup> cells were localized to the active areas of endometrial regeneration near the placental detachment sites at PPD1 with a pronounced expression in the sub-epithelial stromal cells. As regeneration progressed, we continued to observe ALDH1A1<sup>+</sup> cells in the stromal compartment within the placental detachment sites at PPD3 and PPD5, with a progressive, but not statistically significant, increase in ALDH1A1<sup>+</sup> epithelial cells. Collectively, our data demonstrate that ALDH1A1<sup>+</sup> lineage cells participate in the restoration of endometrial architecture and functional compartments in the postpartum phase, even if their direct contribution is transient. Future detailed and mechanistic studies will be necessary to fully characterize their role in this process and their long-term consequence in postpartum regeneration.”

      (5) For Figure 7F, it should be clearly stated in the main text that the results are from one patient sample and the data presented are experimental replicates, so as not to be confused with biological replicates (the same for Supplementary Figure S4). Were B and G in Figure 7 also from one patient?

      Thanks for pointing this out. We have edited the figure legends in the main text and supplemental figures to indicate this.

      Lines 336-337: “…main figures show representative results from one patient sample performed in technical replicates, with additional patient samples included in the supplement…”

      (6) Lines 425-427: "Ovariectomized mice treated with 90-day E2 pellets, on the other hand, showed a complete restriction of ALDH1A1 to the glandular crypts." In Figure 2 S' ALDH1A1+ cells are visible in the LE (the staining is lighter than in the GE but looks real), contradicting this statement.

      This is an important distinction. We have now edited this part of the manuscript to state:

      Lines 458-461: “Ovariectomized mice treated with 90-day E2 pellets, on the other hand, showed enriched ALDH1A1 in the glandular crypts with weak luminal epithelial staining, while the ovariectomized controls had strong ALDH1A1 expression throughout the luminal and glandular epithelium.”

      (7) Lines 466-467: "In cycling mice, we found sporadic cells that expressed both stromal and epithelial markers in the ALDHA1+ cells." These data are not presented.

      We apologize for the confusion, this sentence has been removed from the discussion.

      (8) These data support the role of Aldh1a1+ cells in endometrial epithelial development, but conclusions about their role in repair/regeneration should be tempered as the data are much weaker here.

      We thank the reviewer for their overall assessment. To address this point, we have thoroughly edited the appropriate areas to temper the conclusions and ensure that they are strongly supported by our data. We have also edited the manuscript’s title to reflect this.

      Reviewer #3 (Public review):

      Summary:

      Tan et al demonstrated the importance of ALDH-high cells in the epithelial development in the mouse endometrium, and these cells displayed properties of stem cells.

      We thank the reviewer for their assessment of our manuscript.

      Strengths:

      The findings are solid, supported and validated through a combination of technical methods. I appreciated this combined use of mouse and human endometrial cells to strengthen the findings. Genomic results from a single-cell sequencing dataset were informative as they depicted the different stages of the estrus cycle during the regeneration process. Verification with immunostainings with various markers made it convincing for readers to visualize the cell's location, progression, and status at different timepoints. Utilizing human endometrial cells further demonstrated that the phenomenon observed in mice can be translated to humans.

      This work will greatly advance the understanding of endometrial regeneration for reproductive biologists.

      We thank the reviewer for their expert assessment and positive comments regarding our manuscript.

      Weaknesses:

      No major weaknesses were identified by this reviewer.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) As this study evaluated Aldh1a1+ cells in both the epithelium and stroma, it is recommended that the title and abstract be revised to reflect this.

      Thank you. Both the title and abstract title have been updated to reflect this comment.

      (2) Lines 46-47 in the abstract: "Aldh1a1+ cells expanded during postnatal development, estrus cycling, and following post-partum repair." It is recommended to clarify stromal vs epithelial Aldh1a1+ cell expansion, as only the stromal cells expanded during the estrous cycle. Also, RFP+ epithelial cells were not quantified during postpartum repair and visually appear to decrease (see comment below regarding Figure 5), so this statement is misleading.

      The abstract was edited following the suggestions so that it depicts our results. Similarly, we have addressed the comment regarding Figure 5 and the interpretation of the post-partum regeneration experiments (see comment above for the full explanation and edits to the manuscript).

      (3) Lines 65-69, 186-187: the authors should be clear when describing markers of putative epithelial vs mesenchymal stem/progenitor cells in the introduction (e.g., CDH2/SSEA1/SOX9 for epithelial and SUSD2 for mesenchymal).

      Thank you for this suggestion. The Introduction section now states the following:

      Lines 70-73: “These markers identify specific cell types in the endometrium, with CDH2<sup>+</sup>, SSEA1<sup>+</sup>, and SOX9<sup>+</sup> cells corresponding to endometrial epithelial cells, while SUSD2<sup>+</sup> cells corresponding to mesenchymal endometrial cells enriched in the perivascular regions (Cousins et al., 2021).”

      (4) Lines 78-80: "Studies tracing the fate, ablation, and proliferative capacity of Lgr5+ cells in the uterus identified an Lgr5+ niche that is enriched in the crypts of the glandular epithelium and promotes endometrial regeneration (Seishima et al., 2019)." This statement is incorrect regarding regeneration, as Lgr5 marks stem/progenitor cells in the developing uterus but not the adult, during which endometrial regeneration occurs. Please revise.

      Thank for you for this clarification. The statement has been revised and now reads as follows:

      Lines 82-84: “Studies tracing the fate, ablation, and proliferative capacity of Lgr5<sup>+</sup> cells in the uterus identified an Lgr5<sup>+</sup> niche that marked stem/progenitor cells in the developing uterus (Seishima et al., 2019).”

      (5) Recommend using free-form shapes to outline GE, LE, and EpSC in Fig 2C and stating in the text which clusters correspond to LE and GE (lines 167-169).

      Thank you, this has been edited in Figure 2C and in the text, which now reads as follows:

      Lines 176-177: “Clusters 5, 7, 21 and 16 were classified as glandular epithelial cells, and clusters 0, 2, 3, 10, 14, and 24 were classified as luminal epithelial cells (Figure 2C).”

      (6) "Estrus" refers to the specific stage of the "estrous" cycle. Estrus cycle is incorrect and should be estrous cycle.

      Thank you for pointing this out. It has been corrected throughout the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Suggest increasing the font size for some of the labels in the figures.

      Thank you, we have increased font size in the figures to improve the quality.

      (2) Need to include more details of the human endometrial tissue used in this study: pathology, age, and stage of menstrual cycle.

      We thank the reviewer for the helpful suggestion. The details that are available to us have been included in Supplementary Table S5.

      (3) Line 65-67 - Different endometrial stem cell subsets - SUSD2+ reside in perivascular regions, while other markers are located in glandular epithelium, need revision.

      Thank you, this has been revised in the Introduction. The area now reads as follows:

      Lines 70-73: These markers identify specific cell types in the endometrium, with CDH2<sup>+</sup>, SSEA1<sup>+</sup>, and SOX9<sup>+</sup> cells corresponding to endometrial epithelial cells, while SUSD2<sup>+</sup> cells corresponding to mesenchymal endometrial cells enriched in the perivascular regions (Cousins et al., 2021).

      (4) Include a discussion about the interpretation of their current finding in relation to the dynamic regeneration observed in human endometrium due to menstrual bleeding/tissue breakdown, compared to the cycles of growth and regression that occur in mice.

      This is a great suggestion. We have added the following statement to our Discussion section:

      Lines 537-544: Additionally, our studies in human endometrium extend our characterization of ALDH1A1 as an adult endometrial stem cell marker and emphasize the importance of ALDH1A1+ in the regenerative potential of the endometrium. The conserved hormonal responses between human and mouse endometrium support the hypothesis that cycles of proliferation, differentiation, and regression, whether through resorption/autophagy in mice or menstrual breakdown in humans, are governed by concerted growth factor signaling and dedicated stem cell populations with the capacity to expand and differentiate across repeated cycles of repair. Our detailed studies in both mouse and human models indicate that ALDH1A1+ cells represent a dedicated cell type within the endometrium with the potential to drive repair during adulthood. Collectively, these findings advance our understanding of the mechanisms that control endometrial cycling and regeneration throughout the reproductive lifespan.

      Reference

      Ang, C.J., Skokan, T.D., and McKinley, K.L. (2023). Mechanisms of Regeneration and Fibrosis in the Endometrium. Annu Rev Cell Dev Biol 39, 197-221.

      Bi, W.R., Jin, C.X., Xu, G.T., and Yang, C.Q. (2012). Bone morphogenetic protein-7 regulates Snail signaling in carbon tetrachloride-induced fibrosis in the rat liver. Exp Ther Med 4, 1022-1026.

      Chen, M.Y., Zhao, F.L., Chu, W.L., Bai, M.R., and Zhang, D.M. (2023). A review of tamoxifen administration regimen optimization for Cre/loxp system in mouse bone study. Biomed Pharmacother 165, 115045.

      Cousins, F.L., Murray, A., Esnal, A., Gibson, D.A., Critchley, H.O., and Saunders, P.T. (2014). Evidence from a mouse model that epithelial cell migration and mesenchymal-epithelial transition contribute to rapid restoration of uterine tissue integrity during menstruation. PLoS One 9, e86378.

      Cousins, F.L., Pandoy, R., Jin, S., and Gargett, C.E. (2021). The Elusive Endometrial Epithelial Stem/Progenitor Cells. Front Cell Dev Biol 9, 640319.

      Darzi, S., Werkmeister, J.A., Deane, J.A., and Gargett, C.E. (2016). Identification and Characterization of Human Endometrial Mesenchymal Stem/Stromal Cells and Their Potential for Cellular Therapy. Stem Cells Transl Med 5, 1127-1132.

      Ghosh, A., Syed, S.M., Kumar, M., Carpenter, T.J., Teixeira, J.M., Houairia, N., Negi, S., and Tanwar, P.S. (2020). In Vivo Cell Fate Tracing Provides No Evidence for Mesenchymal to Epithelial Transition in Adult Fallopian Tube and Uterus. Cell Rep 31, 107631.

      Huang, C.C., Orvis, G.D., Wang, Y., and Behringer, R.R. (2012). Stromal-to-epithelial transition during postpartum endometrial regeneration. PLoS One 7, e44285.

      Khanmohammadi, M., Mukherjee, S., Darzi, S., Paul, K., Werkmeister, J.A., Cousins, F.L., and Gargett, C.E. (2021). Identification and characterisation of maternal perivascular SUSD2(+) placental mesenchymal stem/stromal cells. Cell Tissue Res 385, 803-815.

      Kirkwood, P.M., Gibson, D.A., Shaw, I., Dobie, R., Kelepouri, O., Henderson, N.C., and Saunders, P.T.K. (2022). Single-cell RNA sequencing and lineage tracing confirm mesenchyme to epithelial transformation (MET) contributes to repair of the endometrium at menstruation. Elife 11.

      Li, S.Y., Whiteside, S., Li, B., Sun, X., and DeFalco, T. (2025). Mesenchymal-to-epithelial transition of perivascular cells contributes to endometrial re-epithelialization. Nat Commun 16, 10174.

      Niayesh-Mehr, R., Kalantar, M., Bontempi, G., Montaldo, C., Ebrahimi, S., Allameh, A., Babaei, G., Seif, F., and Strippoli, R. (2024). The role of epithelial-mesenchymal transition in pulmonary fibrosis: lessons from idiopathic pulmonary fibrosis and COVID-19. Cell Commun Signal 22, 542.

      Patterson, A.L., Zhang, L., Arango, N.A., Teixeira, J., and Pru, J.K. (2013). Mesenchymal-to-epithelial transition contributes to endometrial regeneration following natural and artificial decidualization. Stem Cells Dev 22, 964-974.

      Pimeisl, I.M., Tanriver, Y., Daza, R.A., Vauti, F., Hevner, R.F., Arnold, H.H., and Arnold, S.J. (2013). Generation and characterization of a tamoxifen-inducible Eomes(CreER) mouse line. Genesis 51, 725-733.

      Rios, A.C., Fu, N.Y., Cursons, J., Lindeman, G.J., and Visvader, J.E. (2016). The complexities and caveats of lineage tracing in the mammary gland. Breast Cancer Res 18, 116.

      Seishima, R., Leung, C., Yada, S., Murad, K.B.A., Tan, L.T., Hajamohideen, A., Tan, S.H., Itoh, H., Murakami, K., Ishida, Y., et al. (2019). Neonatal Wnt-dependent Lgr5 positive stem cells are essential for uterine gland development. Nat Commun 10, 5378.

      Zeisberg, M., Shah, A.A., and Kalluri, R. (2005). Bone morphogenic protein-7 induces mesenchymal to epithelial transition in adult renal fibroblasts and facilitates regeneration of injured kidney. J Biol Chem 280, 8094-8100.

    1. eLife Assessment

      The authors use solid high resolution microscopy techniques to present a model in which SARS-CoV-2 attachment and endocytosis are mediated by heparan sulfate, whereas ACE2 only functions downstream of these early processes. These findings serve as a potentially valuable starting point to examine SARS-CoV-2 entry models that challenge current paradigms. However, examination of the model in a clean heparan sulfate deficient background and in the context of TMPRSS2-dependent plasma membrane fusion are lacking. Thus, the broader impact remains uncertain in the absence of additional controls and orthogonal approaches.

    2. Reviewer #1 (Public review):

      This revised paper investigates how heparan sulfate (HS) engagement functions in the cellular entry of SARS-CoV-2. The authors used a series of microscopy techniques, labeled pseudoviruses and authentic SARS-CoV-2 strains, and cells lacking or expressing HS and/or hACE2 to re-examine the specific stage(s) HS and hACE2 function in the entry process. They suggest that HS mediates SARS-CoV-2 cell-surface attachment and endocytosis, and that hACE2 functions downstream of this to facilitate productive infection. Their results also suggest that SARS-CoV-2 binds clusters of HS molecules projecting 60-410 nm, which act as docking sites for viral attachment. The authors conclude their work establishes a revised entry paradigm in which HS clusters mediate SARS-CoV-2 attachment and endocytosis, with ACE2 acting at some stage downstream. They speculate this idea might apply broadly to other viruses known to engage HS and has translational implications for developing antiviral agents that target HS interactions.

      The strengths of the study include the use of multiple high-resolution microscopy modalities, the tracking of labelled viruses, the use of both pseudoviruses and authentic SARS-CoV-2, and use of primary airway cells. While some studies were performed in the revision to address the Reviewer concerns, which improved the paper clarity, others were cursorily addressed, which limit the impact of the studies. Particularly. experiments were not performed to account for TMPRSS2 expression and plasma membrane fusion. Moreover, addition of studies in which hACE2 is expressed in cells genetically lacking HS were not designed. Thus, it the picture remains unclear picture exactly where downstream hACE2 functions and how this might differ given new structural models of TMPRSS2 activation (PMID: 42050172), which occur after ACE2 recognition of spike on the cell surface.

    3. Reviewer #2 (Public review):

      In the manuscript by Han et al, the authors assess binding of SARS-CoV-2 to heparan sulfate clusters via advanced light microscopy of viral particles. The authors claim that SARS-CoV-2 spike (on the context of pseudovirus and in authentic virus) engages heparan sulfate clusters on the cell surface which then promotes endocytosis and subsequent infection. The finding that HSPGs are important for SARS-CoV-2 entry in some cell types is well described but the authors here attempt to make the claim that HS represents an alternative "receptor" and that HS engagement is far more important than the field appreciates. The data itself appears of appropriate quality and would be of interest to the field, but the overly generalized conclusions lack adequate experimental support. This significantly diminishes enthusiasm for this manuscript as written. Additional controls would be of great benefit.

      Further, it is this reviewers opinion that the findings do not represent a novel paradigm as claimed. HS has been well described for SARS-CoV-2 and other viruses to serve as attachment factors to promote initial virus attachment. A more balanced and nuanced view of their interesting data would be of value.

      Major:

      The authors need to rigorously define a "receptor." This reviewer would argue that a receptor is a host factor that is necessary and sufficient for active promotion of viral entry (genome release into the cytoplasm) while an attachment factor is a host factor that enhances initial viral attachment/endocytosis but is not necessary nor sufficient. The evidence does NOT implicate HS as a receptor under this definition. This is proven in Fig 1 (and elsewhere) in which ACE2 is absolutely required for viral entry.

      The authors should genetically perturb HS biosynthesis in their key assays to demonstrate necessity. HS biosynthesis genes have been shown to be important for SARS-CoV-2 entry into some cells but not others (Huh7.5 cells PMID 33306959 but not in Vero cells PMID 33147444, Calu3 cells 35879413, A549 cells 33574281, and others 36597481. This is inconsistent with the claim that HS is broadly important (beyond the BHK cells overexpressing ACE2 that are used here).

      Is targeting HS really a compelling anti-viral strategy? The data show a ~5-fold reduction. The strengths and limitations of HS targeting should be presented in a more balanced discussion. Animal data showing anti-viral activity of PIX is warranted. This would enhance this claim and also provide key evidence of a relevant role for HS in a more physiologic model.

      The authors provide inadequate discussion into the fact that these studies rely exclusively on cell lines (which also happen to be TMPRSS2 deficient). The role of proteases in the role of HS should be tested in the cell lines and primary cells used as protease expression is a key determinant of the site of fusion.

      An alternative method to disrupt HS (other than PIX) is needed in primary airway cells. A genetic approach would be much more convincing. The authors should also demonstrate whether entry in their primary cell assays are TMPRSS2 vs Cathepsin L dependent (using E64d and camostat for instance) as mentioned above.

      Each figure legend should clearly state how many independent experiments and replicates per experiment were performed.

      All bar plots should show individual dots (i.e. Fig 1G) to better reveal the variance of each dataset.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This paper investigates how heparan sulfate (HS) engagement functions in the cellular entry of SARS-CoV-2. A prevailing model that has been developed over the last five years by work from many laboratories using a variety of biochemical, structural, and microscopic approaches is that HS acts a co-receptor for SARS-CoV-2; its binding to SARS-CoV-2 both concentrates virus on the surface of target cells and allosterically alters the spike protein to promote an "up/open" RBD conformation that enables engagement of the proteinaceous receptor human ACE2 on the cell surface (PMID: 32970989, 35926454, 38055954, 39401361, 40548749). These two events enable plasma membrane fusion (after a cleavage event promoted by plasma membrane TMPSS2) or endocytosis and subsequent pH-dependent fusion (which requires a cathepsin L-mediated cleavage of the spike).

      The authors in this study used a series of microscopy techniques, labeled pseudoviruses and authentic SARS-CoV-2 strains, and cells lacking or expressing HS and/or hACE2 to re-examine the specific stage(s) HS and hACE2 function in the entry process. They suggest that HS mediates SARS-CoV-2 cell-surface attachment and endocytosis, and that hACE2 functions "downstream" of this to facilitate productive infection. Their results also suggest that SARS-CoV-2 binds clusters of HS molecules projecting 60-410 nm, which act as docking sites for viral attachment. Blocking HS binding with pixantrone, a drug under clinical evaluation for cancer (due to its anti-topoisomerase II activity), inhibited SARS-CoV-2 Omicron JN.1 variant from attaching to and infecting human airway cells. The authors conclude that their work establishes a revised entry paradigm in which HS clusters mediate SARS-CoV-2 attachment and endocytosis, with ACE2 acting at some stage downstream. They speculate this idea might apply broadly to other viruses known to engage HS and has translational implications for developing antiviral agents that target HS interactions.

      The strengths of the interesting and technically well-executed study include the use of multiple high-resolution microscopy modalities, the tracking of labelled viruses, the use of both pseudoviruses and authentic SARS-CoV-2, and the use of primary airway cells. Nonetheless, there are issues that need to be addressed to buttress the proposed model compared to earlier ones. These include: (a) the distinction between macropinocytosis and receptor-mediated endocytosis and what this might mean for productive SARS-CoV-2 infection; (b) the need to account for TMPRSS2 expression and plasma membrane fusion; (c) addition of genetic studies in which hACE2 is expressed in cells lacking HS; (d) an unclear picture of exactly where downstream hACE2 functions; and (e) and a need for comparative/additional study of earlier SARS-CoV-2 variants, which preferentially fuse at the plasma membrane.

      We thank the reviewer for the strong support of this manuscript. We addressed the reviewer’s concerns in the Recommendations to the authors. We did not distinguish whether the endocytic route is macropinocytosis or receptor-mediated endocytosis, because it is a separate study beyond the scope of the present work. We did not examine earlier SARS-CoV-2 variants because we considered it a study beyond the scope of the present work, but a good idea that we may work on in the future. For detail on how we address the remaining concerns, please see our response to the reviewer’s Recommendations for the authors.

      Reviewer #2 (Public review):

      In this manuscript by Han et al, the authors assess the binding of SARS-CoV-2 to heparan sulfate clusters via advanced light microscopy of viral particles. The authors claim that the SARS-CoV-2 spike (in the context of pseudovirus and in authentic virus) engages heparan sulfate clusters on the cell surface, which then promotes endocytosis and subsequent infection. The finding that HSPGs are important for SARS-CoV-2 entry in some cell types is well-described, but the authors attempt to make the claim here that HS represents an alternative "receptor" and that HS engagement is far more important than the field appreciates. The data itself appears to be of appropriate quality and would be of interest to the field, but the overly generalized conclusions lack adequate experimental support. This significantly diminishes enthusiasm for this manuscript as written. The manuscript is imprecise and far overstates the actual findings shown by the data. Additional controls would be of great benefit.

      Further, it is this reviewer's opinion that the findings do not represent a novel paradigm as claimed. HS has been well described for SARS-CoV-2 and other viruses to serve as attachment factors to promote initial virus attachment. While the manuscript provides new insight into the details of this process, the manuscript attempts to oversell this finding by applying new words rather than new molecular details. The authors would be better served by presenting a more balanced and nuanced view of their interesting data. In this reviewer's opinion, the salesmanship significantly detracts from the data and manuscript.

      We thank the reviewer for pointing out that our manuscript is of interest to the field. However, we do not think that we oversell our data. hACE2 has been widely considered the receptor (or the binding partner) that mediates SARS-CoV-2 cell-surface attachment, whereas HS is considered only an attachment factor that facilitates SARS-CoV-2 binding with hACE2 at the cell surface. In the present work, we found that HS, but not hACE2, is the cell-surface attachment receptor (or binding partner), whereas hACE2 is not essential for attachment, but acts downstream of virus endocytosis to facilitate viral genome expression. This finding suggests significant modification of the current model by replacing the attachment receptor (or binding partner) from hACE2 to HS, treating HS as a primary receptor rather than an attachment factor, and relocating the hACE2 action site from the cell surface to the endosome. For these reasons, we do not consider these statements overselling our data. However, as the reviewer suggested in his/her specific comments, we revised the manuscript to ensure that we did not overgeneralize our findings (see our responses to the reviewer’s Recommendations to the authors).

      Major Comments:

      The authors need to rigorously define a "receptor" vs an "attachment factor." They also should avoid ambiguous terms such as "receptor underlying ...attachment" and "attachment receptor" (or at least clearly define them). Much of their argument hinges on the specific definition of these terms. This reviewer would argue that a receptor is a host factor that is necessary and sufficient for active promotion of viral entry (genome release into the cytoplasm), while an attachment factor is a host factor that enhances initial viral attachment/endocytosis but is neither necessary nor sufficient. The evidence does NOT implicate HS as a receptor under this fairly textbook definition. This is proven in Figure 1 (and elsewhere) in which ACE2 is absolutely required for viral entry.

      The authors should genetically perturb HS biosynthesis in their key assays to demonstrate necessity. HS biosynthesis genes have been shown to be important for SARS-CoV-2 entry into some cells but not others (Huh7.5 cells PMID 33306959, but not in Vero cells PMID 33147444, Calu3 cells 35879413, A549 cells 33574281, and others 36597481. The authors need to discuss this important information and reconcile it with their data and model if they want to claim that HS is broadly important.

      Is targeting HS really a compelling anti-viral strategy? The data show a ~5-fold reduction, which likely won't excite a drug company. The strengths and limitations of HS targeting should be presented in a more balanced discussion. Animal data showing anti-viral activity of PIX is warranted. This would enhance this claim and also provide key evidence of a relevant role for HS in a more physiologic model.

      The authors provide little discussion of the fact that these studies rely exclusively on cell lines (which also happen to be TMPRSS2-deficient). The role of proteases in the role of HS should be tested in the cell lines and primary cells used, as protease expression is a key determinant of the site of fusion.

      The claim that "SARS-CoV2 JN.1 variant binds to heparan sulfate, not hACE2, in primary human airway cells" is extraordinary and thus requires extraordinary evidence.

      First, PIX reduces attachment by 5-fold, which is not the same as "nearly abolished." Also, anti-ACE2 "nearly abolished" entry in 7D, while PIX did not. If the authors want to make these claims, an alternative method to disrupt HS (other than PIX) is needed in primary airway cells. A genetic approach would be much more convincing. The authors should also demonstrate whether entry in their primary cell assays is TMPRSS2 vs Cathepsin L dependent (using E64d and camostat, for instance) as mentioned above.

      Each figure should clearly state how many independent experiments and replicates per experiment were performed. What does "3 experiments" mean? Are these three independent experiments or three wells on one day?

      In the well-accepted current model, hACE2 is considered the receptor mediating SARS-CoV-2 cell-surface attachment, entry into cells, and infection, whereas HS is an attachment factor that facilitates SARS-CoV-2 binding to hACE2 at the cell surface. The present work revises this view: HS is the SARS-CoV-2 attachment receptor mediating virus docking at the cell surface, with ACE2 acting downstream of virus endocytosis to enable SARS-CoV-2 infection in the cell types examined.

      We made this point clearer throughout the newly revised manuscript. We define the attachment receptor as the docking site where the virus binds to the cell surface. We directly showed with several super-resolution imaging techniques that the virus docks at HS clusters.

      The cited CRISPR-screen literature supports context-dependent host-factor usage. However, the absence of HS biosynthesis genes from a given screen does not prove that HS is irrelevant in that cell type; it only indicates that HS biosynthesis was not detected as a genetic dependency under that assay’s conditions. Such negative results can reflect screen sensitivity, incomplete knockout, pathway redundancy, or viral dose/stringency. In the revised manuscript, we included the following in the Discussion:

      “While some studies using genome-wide CRISPR screening to identify genes involved in SARS-CoV-2 reveal genes for HS biosynthesis, others do not (45-50). The negative result, which might reflect screen sensitivity, incomplete knockout, pathway redundancy, or viral dose/stringency, needs to be verified with specific gene knockout.”

      The ~5-fold reduction is likely due to the inhibitor not completely abolishing HS-virus binding. We revised the Discussion to strengthen the suggestion that targeting the virus cell-surface attachment by interfering HS binding is a therapeutic strategy to prevent and treat COVID-19, as in the following:

      “Interfering with HS binding has been suggested as a therapeutic strategy to prevent and treat many viral infections that depend on HS for entry, including COVID-19 [1, 2, 9, 12]. Supporting this strategy, disrupting Spike–HS interactions, including inhibition by heparin and related glycans, reduces SARS-CoV-2 attachment/entry [51]. Clinical evaluation of inhaled/nebulized unfractionated heparin has reported improved clinical outcomes without major bleeding signals, supporting the feasibility of targeting airway-surface HS interactions [52]. HS mimetics, such as pixatimod (PG545), inhibit SARS-CoV-2 infection and exhibit greater potency than heparin in assays measuring inhibition of Spike/ACE2 engagement and viral infectivity [53]. These reports support the translational potential of therapeutically interfering with virion–HS binding. However, this strategy has not been the focus for developing methods to prevent and treat COVID-19, likely because HS is considered only a regulator that is not essential for SARS-COV-2 entry. Our finding that HS is the attachment receptor re-emphasizes the importance of perturbing virus-HS binding, the first step of the viral entry, to efficiently block SARS-CoV-2 infection. Further supporting this view, inhibition of HS binding with a clinically used HS-binding agent, pixantrone, inhibits authentic SARS-CoV-2 JN.1 subvariant binding with HS on the cell surface and infection in primary human airway cells (Figs. 6, 7). These results suggest a combinatorial anti-SARS-CoV-2 strategy: early HS blockade to prevent attachment combined with ACE2 targeting to inhibit post-attachment steps”

      We include a sentence in the Discussion that our suggestions are limited to the cells we examined as below.

      “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      Three experiments refer to three independent experiments. We added “independent” accordingly throughout the manuscript.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors define a new paradigm for the attachment and endocytosis of SARS-CoV-2 in which cell surface heparan sulfate (HS) is the primary receptor, with ACE2 having a downstream role within endocytic vesicles. This has implications for the importance of targeting virion-HS interactions as a therapeutic strategy.

      Strengths:

      The authors show that viruses are internalized via dynamin-dependent endocytosis and that endocytic internalization is the major pathway for pseudotyped SARS-CoV-2 genome expression. They show that HS-mediated viral attachment is a critical step preceding viral endocytosis and also subsequent genome expression. Further, they show that hACE2 acts downstream of endocytosis to promote viral infection, and may be co-internalised with virions after HS attachment. Pseudotyped virus and authentic SARS-CoV-2 provide similar results. In addition, the authors demonstrate that remarkable clusters of multiple HS chains exist on the cell surface, visualised by a number of elegant microscopy methods, and that these represent the docking sites for virions. These visualisations are an important general contribution in themselves to understanding the nanoscale interactions of HS at the cell surface.

      The use of a complementary range of methods, virus constructs, and cell models is a strength, and the results clearly support the conclusions.

      Overall, the results convincingly demonstrate a different model to the currently accepted mechanism in which the ACE2 protein is regarded as the cell surface receptor for SARS-CoV-2. Here, the authors provide compelling evidence that cell surface clusters of HS are the primary docking site, with ACE2 interactions occurring later, after endocytosis (whilst still being essential for viral genome expression). This is an exciting and important landmark evidence which supports the view that HS-virion interactions should be viewed as a key site for anti-viral drug targeting, likely in strategies that also target the downstream ACE2-based mechanism of viral entry within endosomes.

      We thank the reviewer for the strong support of the present work.

      Weaknesses:

      This reviewer identified only minor points regarding citing and discussing other studies and typos, which can be corrected.

      We have addressed these points in the revised manuscript. For detail, please see our response to the reviewer’s Recommendations to the authors.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Pathway of internalization.

      The authors show clearly that labeled SARS-CoV-2 (pseudovirus or authentic virus) can become internalized in cells lacking hACE2, and this process depends on HS. However, they also show that this pathway is non-productive with regard to infection. Are the entry vesicles mediated by HS alone, HS + hACE2, and hACE2 alone the same? Or does the combination of co-receptor (HS + hACE2) drive SARS-CoV-2 into endocytic vesicles, whereas HS alone promotes macro- or micropinocytosis (lines 361-362). If HS alone directed SARS-CoV-2 into a non-productive entry vesicle, then hACE2 likely would be acting concurrently with HS and not downstream. A more detailed analysis of the different entry vesicles/pathways that occur with HS alone, HS + hACE2, and hACE2 alone is needed.

      During endocytosis, we did not detect a difference in the size distribution of virus-containing vesicles between BHK (HS alone) and BHK<sub>hACE2</sub> cells (HS+hACE2) (Fig. 2D). The similarity in the vesicle size suggests a similar endocytic path with HS alone or with HS + hACE2. In the revised manuscript, we added the following sentence.

      “Third, 3D-STED imaging showed that A490-labeled vesicle’s full-width-at-half-maximum (W<sub>H</sub>) was 363 ± 17 nm (n = 55) in BHK cells, similar to that (333 ± 13 nm, n = 70) in BHK<sub>hACE2</sub> cells (Fig. 2B-D), supporting a similar endocytic path regardless of hACE2 presence or not.”

      (2) TMPRSS2 and plasma membrane fusion.

      Although the authors allude to membrane fusion as an alternate mechanism of entry, their mechanistic experiments do not address the roles of HS and hACE2 in this process, possibly because their BHK and other cells do not co-express significant levels of TMPRSS2. While many Omicron variants preferentially enter cells via endocytosis (relative to antecedent strains in the pandemic) because of spike mutations that reduce cleavage by TMPRSS2 (PMID: 35104837, 36625591, 35145066), plasma membrane fusion can still occur. The authors should add experiments with co-expression of TMPRSS2/hACE2 [with or without HS] and earlier SARS-CoV-2 variants to establish the role of HS in plasma membrane fusion. Also, are there differences in entry pathways if viruses are prepared in cells expressing TMPRSS2?

      We thank the reviewer for this important comment and agree that our mechanistic experiments were not designed to address TMPRSS2-supported plasma membrane fusion. The reviewer’s suggestion for direct testing of HS function in TMPRSS2-supported plasma membrane fusion, including hACE2/TMPRSS2 co-expression and comparison with earlier SARS-CoV-2 variants, will require dedicated experiments and detection of the fusion pathway that we have not yet designed. It is beyond the scope of the present work. In the revised manuscript, we clarify that the observed ACE2-independent uptake and the predominance of endocytic entry refer to the tested cell systems and do not exclude TMPRSS2-dependent plasma membrane fusion in other cell types, as in the following.

      “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      (3) Experiments with hACE2 in cells lacking HS.

      Apart from drug treatment (heparinases or pixantrone) studies shown, the current studies do not directly address whether expression of hACE2 on human cells can allow for endocytosis and productive infection in the complete and genetic absence of HS. The only experiments that use genetically deficient cells are the CHO [hamster] cell studies, and these cells lack hACE2 expression. The authors should knock out a key HS biosynthesis gene (e.g., B4GALT7) in more relevant human cells (e.g., A549-hACE2; ideally with sorted subpopulations having different levels of surface hACE2 expression) and assess endocytosis and infection. This is important given studies in the literature by others suggesting that KO of HS expression reduces but does not abrogate SARS-CoV-2 infection.

      We thank the reviewer for these comments. We showed that virus endocytosis is independent of hACE2 (Fig. 1). The reviewer’s question is whether hACE2 alone can allow for endocytosis of viruses. We have shown that in either BHK (without hACE2) or BHK<sub>hACE2</sub> cells (BHK cells expressed with hACE2), heparinase I/II/III mixture (HPRase) nearly abolished cell-surface immunolabelled HS (Fig. 3D), reduced cell-surface virus attachment by ~83-85% (Fig. 3E), reduced viral uptake by ~80% (Fig. 3F, 3G). These results suggest that hACE2 is not essential for viral attachment and endocytosis. We did not test whether hACE2 alone (without HS) plays a minor role for viral attachment and endocytosis, because to our knowledge, HS is present in nearly every cell. Under this physiological condition, it is HS, not hACE2, that plays an essential role in viral cell-surface attachment and endocytosis. In the revised manuscript, we added a sentence admitting that we did not test whether hACE2 alone is sufficient to support viral uptake and productive infection, as in the following.

      “Our data suggest that, under physiological conditions in which HS is present at the cell surface, hACE2 is not essential for viral cell-surface attachment or endocytosis. We do not know whether hACE2 expression alone, in the absence of HS, can support viral cell-surface attachment and endocytosis.

      (4) hACE2 function in entry.

      In many places, the authors suggest that hACE2-spike functional interaction occurs "downstream" of HS-dependent binding and endocytosis (e.g., lines 25, 33, 48, 210, 309, 312, 318, 333, 346). However, in their model, it is not clear where exactly this interaction occurs. Are the authors suggesting that this spike binds hACE2 on the cell surface, but this has nothing to do with endocytosis, or that the interaction with hACE2 is occurring at a post-entry step? Can they experimentally demonstrate the stage at which hACE2 is functioning? Is it the same or different in cells lacking HS? What about when TMPRSS2 is present?

      We showed that viral attachment and endocytosis are independent of hACE2, whereas entry as determined by viral gene expression, depends on hACE2. We also showed that most virions bind to HS, not hACE on the cell surface. Based on these results, we propose a model that hACE2 functions downstream of virion endocytosis. We cannot rule out the possibility that a small subset of viruses can also bind to hACE2 after their binding with HS at the cell surface.

      We have not been able to design an experiment to visualize hACE2 mediated virion fusion in endosomes, where hACE2 may facilitate virus fusion and delivery of viral genomes to the cytosol. Productive infection requires only a limited number of successful virion–hACE2 engagement events. While many internalized virions can be visualized, the specific virion or vesicle that ultimately gives rise to productive infection cannot be identified from the present imaging data. This makes it difficult to trace the precise stage or compartment in which the functionally relevant spike–hACE2 interaction occurs. In the revised manuscript, we added a paragraph discussing this limitation as below.

      “Our data suggest that, under physiological conditions in which HS is present at the cell surface, hACE2 is not essential for viral cell-surface attachment or endocytosis. We do not know whether hACE2 expression alone, in the absence of HS, can support viral cell-surface attachment and endocytosis. Although our data suggest that hACE2 functions downstream of endocytosis to facilitate viral fusion at the endosome for genome delivery to the cytosol, we do not know whether hACE2 binding with the virus occurs at the cell surface or endosomes. The binding may occur in both places, but not essential for virus attachment and endocytosis.”

      (5) Other comments.

      (a) Figure 1A. "Antibody" is misspelled.

      Corrected. Thank you.

      (b) The imaging experiments with pseudoviruses and authentic viruses lack any information on the multiplicity of infection or the number of virions added per cell. If this is particularly high and non-physiological (e.g., >100), is it possible that such conditions might enable viruses to enter [dominantly] through secondary [non-infectious] pathways?

      To address the reviewer’s concern, we used flow cytometry to measure cell-associated VSV-S signal as we diluted the virus by ~600-fold. We found that the V-A647 attachment at the cell surface of BHK cells was similar to that in BHK<sub>hACE2</sub> cells over a ~600-fold dilution of the virus (Fig. S7), indicating that the virus cell-surface attachment is independent of hACE2 across a wide range of virus concentrations. In the revised manuscript, we included the following sentence and Fig. S7 (Supplementary Information).

      “Flow cytometry also showed that the V-A647 attachment at the cell surface of BHK cells was similar to that in BHKhACE2 cells over a ~600-fold dilution of the virus concentration (Fig. S7), indicating that the virus cell-surface attachment is independent of hACE2 across a wide range of virus concentrations.”

      (c) Figure 1C and elsewhere. Most of the internalization studies rely on various imaging modalities to demonstrate the pseudovirus or virus on or in the cell. The experiments would be strengthened by inclusion of data from orthogonal binding/internalization assays that measuring virion-associated viral RNA on the surface [4oC binding assay] or inside the cell [after a 37oC temperature shift and exogenous proteinase K and RNAse A treatment]) - such assays can be performed at much lower MOI (e.g., <1, addressed comment #2 above) an also allow more objective quantitation and kinetic analyses of virus internalization (e.g., 0, 5, 15, 30 min at 37oC).

      We demonstrate virion attachment and uptake using multiple approaches, including confocal, STED, and EM analysis, showing virions with the expected morphology at the cell surface and in the cytosol. Furthermore, flow cytometric analysis provides population-level quantitation supporting the same overall conclusion. Thus, while we appreciate and agree that an RNA-based binding/internalization assay would provide additional information, we do not consider it essential to the main conclusion of this work.

      (d) Figure 2. (i) Is there any indication of which vesicles the bath dye is in? Is most of this fluid taken up by micropinocytosis? Are these the same vesicles where the virus that is destined for productive infection (HS/hACE2 engaging) transits? (ii) In all panels, can the authors clearly indicate/label which cells are being used (BHK or BHK-hACE2)? (iii) For the studies with dynasore or dominant-negative dynamin-2-K44A, the readout is at 24 h, a late timepoint, which also could affect virus egress and spread. Can the studies be repeated at much earlier time points (e.g., 15 min to 2 h) to demonstrate that viruses are internalized via dynamin-dependent endocytosis in these cells?

      (i) The bath dye A490 was used as a fluid-phase marker for endocytic uptake, rather than as a marker for a specific vesicle class or intracellular compartment. In principle, any vesicle that takes up extracellular fluid could become labelled by this approach. Since nearly all viruses are in the A490-containing vesicles, productive virus infection must come from some of these vesicles.

      (ii) In the revised Fig. 2 legends, we explicitly indicate which cells are used for each panel.

      (iii) To address the reviewer’s concern, we examined earlier time points for dynasore treatment and found that the virus uptake and genome expression were already markedly reduced at 1 h and 8 h after virus incubation. In the revised manuscript, we described these results as below and in Fig. S5.

      “Fourth, dynasore or dominant-negative dynamin 2-K44A overexpression, which inhibits fission of dynamin-dependent endocytosis [28-30], substantially reduced V-A647 internalized 1-24 h after viral incubation (Figs. 2F-G, S5).

      In addition to inhibiting V-A647 endocytosis, dynasore or dynamin 2-K44A inhibited V-EGFP expression 8-24 h after virus incubation by ~66-77% (Figs. 2F-G, S5), suggesting that endocytosis is the main route for viral genome expression.”

      (e) Line 225. "Envelop" should be "envelope".

      Corrected, thank you.

      (f) Line 235. The authors should clarify that they conclude that the "Omicron variant" of SARS-CoV-2 enters "BHK" cells indistinguishably from VSV-S.

      Thank you for pointing this out. We have rephrased the conclusion as “…omicron variant of SARS-CoV-2 enters BHK cells indistinguishably to VSV-S.”

      (g) Line 278. What happens to virus binding if the authors ectopically express hACE2 in CHO-K1 WT and CHO-pgsA-745 cells?

      We did not perform this experiment (see also our response to major comment 3 above).

      (h) Lines 280-281 and elsewhere (line 635). The authors state "pixantrone (PIX), a drug under clinical trial that binds HS to inhibit HS binding with proteins...." The authors should clarify that the drug is under clinical evaluation for cancer treatment because of its DNA intercalating activity (and not its HS binding activity) and cite any relevant ongoing trials. Also, in line 635, is reference #46 correct?

      As suggested, we modified this sentence as “pixantrone (PIX), a drug under clinical trial for cancer treatment due to its DNA intercalating activity, which can bind HS to inhibit HS binding with proteins”

      (i) Line 281-282. The authors should confirm in a Supplementary Figure that the anti-hACE2 antibody used blocks SARS-CoV-2-JN.1 binding to ACE2.

      In Figure 7D, we showed that PIX and anti-hACE2 antibody block SARS-CoV-2-JN.1 infection, suggesting that anti-hACE2 blocks SARS-CoV-2-JN.1 binding with hACE2.

      (j) Figure 7B. Can hACE2 co-localization be added to this panel?

      We did not perform this experiment. We addressed the role of ACE2 in these airway cells in subsequent panels of Fig. 7.

      (k) Figure 7C. The quantitative data show a 50% reduction in binding signal with pixantrone, whereas the microscopy images appear to show a much greater effect. Can more representative images be shown so that the data better corresponds?

      A ~50% effect is not as visually obvious as the current Fig. 7C. Therefore, we chose not to change the images. However, the statistics in Fig. 7C (right) clearly indicate an average effect of about 50%, as the reviewer pointed out.

      (l) In the Discussion, it is not necessary to use Figure callouts (as done in the Results). Please remove, with the exception of reference to the model.

      We prefer to call out Figures in the Discussion so that we can remind the readers where to find the data. The readers may choose to neglect these callouts. But some readers may read most the discussion part without going through the results carefully. In this case, the figure callouts may help these readers.

      (m) Please delete all references to "new" or "novel" models. It is unnecessary.

      As the reviewer suggested, we deleted “new” and “novel” throughout the revised manuscript.

      (n) Figure legends. Please make sure each panel indicates the # of independent experiments performed. This is included for some but not all panels. Also, a few panels use an unpaired t-test where an ANOVA with multiple comparisons is required (e.g., Figure 1G and S1).

      We agree that, for the three-group sub-comparisons shown within Fig. 1G and Fig. S1, the relevant analyses should account for multiple comparisons. In the revised manuscript, we therefore analyzed these predefined three-group subsets using ordinary one-way ANOVA followed by Dunnett’s multiple-comparisons test, with BHK or Vero used as the reference group as appropriate. The two-group comparisons were analyzed using unpaired two-tailed t-tests.

      Reviewer #2 (Recommendations for the authors):

      (1) It is well established that ACE2 is the receptor for SARS-CoV-2. The authors should not downplay this by saying it is "widely assumed", "typically thought", etc. The specific molecular details at various stages of entry (i.e, the role of HS) remain a bit unclear, but it is disingenuous to imply ACE2 is not the bona fide receptor by any conventional definition.

      The present work does not challenge the well-established view that ACE2 is the receptor for SARS-CoV-2 entry/infection, but suggests that HS is the SARS-CoV-2 attachment receptor mediating virus docking at the cell surface, whereas ACE2 acts downstream of virus endocytosis to enable SARS-CoV-2 infection in the cell types examined. We made this point clearer throughout the revised manuscript. We define the attachment receptor as the docking site where the virus binds to the cell surface. We directly showed with several super-resolution imaging techniques that the virus docked at the HS clusters.

      As the reviewer suggested, we removed “assumed” and “typical” and clarify that our findings do not challenge this concept. For example, we modified the abstract

      “Virus entry is thought to involve binding a unique receptor for cell attachment and cytosolic entry. For SARS-CoV-2 underlying the COVID-19 pandemic, angiotensin-converting enzyme 2 (ACE2) is widely assumed as the receptor.”

      as

      “Virus entry is thought to involve binding a unique receptor for cell attachment and cytosolic entry. For SARS-CoV-2 underlying the COVID-19 pandemic, angiotensin-converting enzyme 2 (ACE2) is widely considered the receptor for cell-surface attachment and subsequent cell entry.”

      (2) When the authors state pseudovirus internalization is independent of ACE2, they should clarify that this is the case in cells not expressing TMPRSS2. Most physiologically relevant cell types express TMPRSS2, which will facilitate entry at the plasma membrane.

      As the reviewer suggested, we included the following sentence in the Discussion section: “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      (3) Line 130: "Endocytic internalization is the main viral infection pathway" and Line 180-181 is not precise and should be rephrased to include the cell types described in the figure. This may be true in BHK-ACE2 cells, but the evidence in this section does not show that this is universally or broadly true.

      We agree and have revised these sentences to limit the conclusions to the experimental context directly supported by our data. Specifically, our results support endocytic uptake as the major route leading to pseudovirus genome expression in the pseudovirus assays and cell types examined here, rather than as a universal entry mechanism for SARS-CoV-2 across cell types. We have therefore modified the subsection title and the relevant sentence in the Results to explicitly refer to the tested cells/assays.

      Across the revised manuscript, we have accordingly revised the text to distinguish initial virion docking/attachment from productive entry, to acknowledge ACE2 as the established receptor for productive infection, and to limit our mechanistic conclusions to the cellular systems directly tested here.

      (4) All bar plots should show individual dots (i.e., Figure 1G) to better reveal the variance of each dataset.

      While we respect the reviewer’s suggestion, this is not required in the journal style. We prefer plotting bar graphs without individual data points, which often makes it difficult to see the mean values.

      (5) Line 57: This is not accurate. HIV uses a receptor and a co-receptor, for instance.

      We thank the reviewer for noting this inaccuracy. We agree that viral entry frequently involves coordinated engagement of multiple host factors rather than a single receptor, for example, HIV requires both a primary receptor and a co-receptor. We have revised the statement in the Introduction (Line 57–58) to reflect that entry can involve receptors together with co-receptors and/or attachment factors, which collectively facilitate membrane fusion or endocytic uptake.

      In the Introduction (Line 57), we replaced the sentence with “Viral entry is often initiated by engagement of host receptors and associated co-factors that together facilitate subsequent viral membrane penetration.”

      (6) Line 60: "most" --> "many"

      As suggested, we have changed “most” to “many”.

      (7) Remove "clinically relevant" in reference JN.1, as JN.1 is not circulating currently. A more appropriate term could be "full-length" or "authentic", or "wild-type".

      As suggested, we changed it to “authentic”.

      Reviewer #3 (Recommendations for the authors):

      (1) The authors omit to mention the work of Zhang et al, 2023 Nature Comms. "Host heparan sulfate promotes ACE2 super-cluster assembly and enhances SARS-CoV-2-associated syncytium formation". These authors also use PIXN and MTN compounds and define different mechanisms based on ACE2 clustering for virus entry. The authors should mention this work in the Discussion and try to reconcile the different findings.

      As suggested, we include the following discussion in the revised manuscript.

      “Consistent with this possibility, HS may promote spike-dependent ACE2 super-cluster assembly at the cell surface and enhance SARS-CoV-2–associated syncytium formation, suggesting that HS may organize ACE2 nanoscale architecture in a cell–cell fusion context [43].”

      (2) The authors should strengthen their case for the validity of HS-virion interactions as a therapeutic target by mentioning studies showing effectiveness of interference with HS-Covid interactions by heparin and other investigational drugs eg. first study to demonstrate heparin inhibition of SARS CoV2 attachment, Mycroft-West et al, Thromb Haemostatis, 2020; recent report of successful clinical trials of nebulized heparin, The Lancet, Sept 2025; and the superior efficacy of HS mimetic Pixatimod/PG545 compared to heparin (Guimond et al 2022 ACS Chemical Sciences).

      We thank the reviewer for this suggestion and add the following paragraph with citations the reviewer mentioned in the Discussion section.

      “Interfering with HS binding has been suggested as a therapeutic strategy to prevent and treat many viral infections that depend on HS for entry, including COVID-19 [1, 2, 9, 12]. Supporting this strategy, disrupting Spike–HS interactions, including inhibition by heparin and related glycans, reduces SARS-CoV-2 attachment/entry [51]. Clinical evaluation of inhaled/nebulized unfractionated heparin has reported improved clinical outcomes without major bleeding signals, supporting the feasibility of targeting airway-surface HS interactions [52]. HS mimetics, such as pixatimod (PG545), inhibit SARS-CoV-2 infection and exhibit greater potency than heparin in assays measuring inhibition of Spike/ACE2 engagement and viral infectivity [53]. These reports support the translational potential of therapeutically interfering with virion–HS binding.”

      (3) Figure 1a: incorrect label for antibody.

      Corrected, thank you.

      (4) Some misspellings noted in the manuscript, e.g., MINFLLUX, so please recheck the manuscript for typos.

      We have rechecked the manuscript and corrected the typos.

    1. eLife Assessment

      This important study investigates the mechanisms by which Mycobacterium tuberculosis suppresses protective Th17 differentiation during infection. Using ESX-1- and PDIM-deficient mutants of M. tuberculosis, which lack functional eccC1 and fadD28 respectively, the authors demonstrate that these virulence factors actively restrict IL-17 responses via a T-bet-dependent pathway, independent of overall bacterial attenuation or infection duration. While the precise molecular interactions by which ESX-1 and PDIM modulate Th17 differentiation warrant further elucidation, the experiments are rigorously designed, the findings are clearly presented, and the evidence is solid for a specific role for these factors in limiting protective IL-17 immunity in M. tuberculosis infection.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript examines the factors that restrict the induction of IL-17-producing T cells during Mycobacterium tuberculosis (Mtb) infection. The authors show that neither infectious route, nor duration of infection are responsible. But they do show that mice that lack the Th1-defining transcription factor, a finding consistent with prior reports in the field of immunology. They also show that 2 highly attenuated Mtb mutants in ESX-1 and PDIM, two well-known Mtb virulence factors, do induce IL-17 producing T cells. In contrast, Mtb mutants in mmpl4 are also similarly attenuated, but do not induce IL-17-producing T cells, suggesting that this property is not simply a result of attenuation but due to specific properties of ESX-1 and PDIM-deficient mutants.

      Strengths:

      (1) It is interesting that mice infected with ESX-1 and PDIM mutants have increased induction of Th17 cells.

      (2) Data is solid and convincing throughout.

      Weaknesses:

      There are two main criticisms:

      (1) B6 mice, compared to humans are known to be very Th1 skewed and the Th1 transcription factor T-bet is known to be a strong inhibitor of Th17 responses. Thus, these Th17 inhibitory factors may be stronger in B6 mice than humans, as many humans do make Th17 responses to Mtb infection.

      (2) The molecular insights about how Th17 induction is somewhat limited. Tbet induction is known to restrict Th17 development and this is a t cell intrinsic mechanism. In contrast, the IL-23 association revealed seems to be extrinsic to T cells and to act on T cells. It is not clear these factors related to each other in restricting Th17 induction.

      Additional points:

      (1) The manuscript states, "Under the conditions where Th17s are highly induced, mice infected with either ΔESX-1 or PDIM lacking Mtb, the Il17a-/- mice had ~3-5 fold higher CFU than WT mice (Figures 3F-G). These results indicate that the induction of Th17s is not dependent on the attenuation of Mtb in general, but instead Mtb utilizes ESX-1 and PDIM to suppress the induction of a Th17 response that enhances protection against Mtb infection." One consideration, however, is that ESX-1, PDIM, and mmpl4 mutants all have similarly reduced CFUs in the lung, but have different CFUs in the lung-draining LN where T cell priming occurs? The bacterial burden in the LN may be more important for regulating T-bet, IL-23, and Th17 differentiation, since the LN is where T cell priming occurs, than the CFU in the lung. Perhaps ESX-1 and PDIM mutants have reduced CFU in the LN, but mmpl4 does not. This difference in LN burdens may be the primary driver of Th17 priming, as high avidity interactions are thought to be an important driver of T-bet induction. Thus, without examining the LN, some questions remain regarding the conclusion that the altered Th17 response in the attenuated strains is not due to the attenuation itself. However, I agree the CFU in the LN probably reflects that in the lung, and if so, the author's conclusions would be sound.

      (2) Do LN cDC1 and high levels of IL-12 p35 manifest in mice infected with the mmpl4 mutant? Likewise do LN cDC2's express low levels of IL-12 p19 (akin to those infected with WT Mtb). If these observations for ESX-1 and PDIM mutants are mechanistically linked to the increased numbers of Th17 cells, then you would expect mice infected with mmpl4 mutants to be more like those infected with WT Mtb than to those infected with ESX-1 and PDIM mutants. These experiments would help provide more convincing evidence that the identified mechanisms are due specifically to outcomes regarding Th17 induction. However, I agree the author's conclusions are the most likely explanation given the current data.

    3. Reviewer #2 (Public review):

      In this manuscript, the authors tackle an important question of why IL-17 production and TH17 responses are lower than expected during Mtb infection. The authors identify an axis of cross-regulation between TH1 and TH17 cells and provide data to support roles for Mtb virulence factors ESX1 and PDIM in promoting TH1 responses and/or suppressing TH17 responses.

      Strengths:

      The strengths include the significance of the work, the combination of host and Mtb genetic models to dissect the mechanistic basis for regulation of IL-17 production from T cells during infection, and the rigor of the experiments. There are a number of exciting findings from the work, including the cross talk between T cell responses and the impact of ESX1 and PDIM on these responses. It is particularly striking that that IL17a deficient mice partially rescue the attenuation of ESX-1 and PDIM mutants.

      Comments on revised version.

      The revised manuscript has tempered a lot of the language in the original text to more accurately state (and not overstate) interpretations of the data. The claim that the effect is independent of route of infection seems a little too large of a claim when only two routes were tested (aerosol and intranasal). And although the authors revised the results section to acknowledge the contribution of an IFNg-dependent suppression of IL-17 production from T cells, the abstract has not been updated and still claims that all effects are independent of IFNg.

    4. Reviewer #3 (Public review):

      Summary:

      The manuscript by Zilinskas et al seeks to understand the mechanisms underlying the ability of Mtb to suppress Th17 differentiation. As Th17 responses are needed for protective immunity against TB, this is an important topic of investigation. They use Mtb mutants that lack eccC1 (from ESX-1 locus) and fadD28 (encoding PDIM) and implicate a Tbet-dependent pathway by which Mtb modulates Th17 differentiation. The mechanism by which ESX-1/PDIM function to impact Th17 differentiation is, however, unclear, which limits the novelty of the results.

      Strengths:

      Understanding how Mtb limits Th17 differentiation has implications for vaccine development. Comparative study of KO mice and Mtb mutants is a strength.

      Weaknesses:

      (1) Addressing several questions related to the Tbet KO mouse experiments would strengthen the study. Do the Tbet KO mice have elevated IL-4/5/13 (which has been previously reported in non-TB studies) in addition to IL-17? The lack of Th17 cells in the IFNg KO compared to the Tbet KO may reflect a difference in timing, since only 3-week data are shown; earlier and later time points would provide a better interpretation. The authors do not present any data on neutrophil infiltration in WT vs Tbet KO vs IFNg KO mice. Since IL-17 is known to be important for recruiting neutrophils to the lung, neutrophil data are important for clarifying the mechanism underlying the CFU outcomes.

      (2) While IL-23 is important for sustaining IL-17 production, IL-6, TGF-b and/or IL-1β are necessary for Th17 polarization. What were the levels of these cytokines in DCs in the lung? (Fig 5). Additionally, Tbet-deficient DCs exhibit impaired activation of antigen-specific Th1 cells and have reduced IL-12 production. Given the data showing higher IL-17 levels in Tbet KO mice, the authors should provide information on the DC phenotype (IL-23, IL-6 etc) in the Tbet KO experiments.

      (3) The mechanism by which ESX-1/PDIM function to impact Th17 differentiation is not clear. While data showing a role for ESX-1 and PDIMs in inhibiting Th17 responses is interesting, there is no insight into the potential mechanism of action. Fig 3 showing reduction in IFNg+ CD4 T cells after infection with eccC1 and fadD28 mutants suggests that this outcome is due to a lower bacterial load relative to WT Mtb at the 3-week time point. Since IFNg is known to suppress IL-17, the higher levels of Th17 cells could be due to the reduction in IFNg due to the attenuated growth of the mutants.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      There are two main criticisms:

      (1) It is not clear how much the factors uncovered here are true beyond B6 mice. B6 mice, compared to humans, are known to be very Th1-skewed, and Tbet is a strong inhibitor of Th17-specific T cells. Many people make IL-17-producing T cells in response to Mtb infection.

      We appreciate the point that not all findings in mice are directly translatable to humans. The B6 mouse is widely used as a model organism for tuberculosis due to its tractability and the wealth of genetic tools available for this strain. While it is true that many individuals do produce Th17 cells after infection with Mtb, humans are still very Th1-dominant, and not all infected individuals produce Th17 cells. We can speculate that the mechanisms outlined in this paper may contribute to the reasons that Th17 responses are not more robust in humans, a finding that may be useful in guiding vaccine design in the future.

      (2) Very few novel insights are mechanistically revealed about how Th17 induction is restricted by Mtb. Tbet induction is known to restrict Th17 development, and this is a T-cell intrinsic mechanism. In contrast, the IL-23 association revealed seems to be extrinsic to T cells and to act on T cells. How, if at all, are these factors related to each other in restricting Th17 induction? Also, the conclusion that it is not a result of attenuation is not completely convincing.

      While it is established that Th1 differentiation can inhibit Th17 differentiation, we believe that rigorously demonstrating this genetically in the context of Mtb infection remains important. Moreover, it cannot be assumed that IL-17 elicited by dampening the Th1 response can lead to enhanced control of infection. We view addressing this as a significant contribution. Furthermore, we also show that the ESX-1 and PDIM virulence factors are functionally linked by suppression of IL-17 responses. The effect is unlikely to be simply due to attenuation of the strains as an equally attenuated control strain does not elicit Th17 cells. We believe that these insights are both novel and important for understanding immune responses to Mtb.

      Other points:

      (1) The authors show that mice infected with a deficiency in ESX-1 have more IL-17-producing CD4 T cells in response to stimulation with an ESAT-6 peptide pool (Figure 3B). Because ESAT-6 is encoded by ESX-1, why do mice infected with this Mtb mutant have any ESAT-6-specific T cells? Is it an incomplete knockdown?

      The ESX-1 knock-out M. tuberculosis Erdman strain is a ΔEccC1 mutant. This strain can produce Esat-6 but cannot secrete Esat-6 out of the bacterial cell. Thus Esat-6 protein is present and able to be processed for MHC-II presentation. We also use Ag85b peptide pool stimulation and report similar effects as Esat-6 peptide pool stimulation.

      (2) The manuscript states, "Under the conditions where Th17s are highly induced, mice infected with either ΔESX-1 or PDIM lacking Mtb, the Il17a-/- mice had ~3-5 fold higher CFU than WT mice (Figures 3F-G). These results indicate that the induction of Th17s is not dependent on the attenuation of Mtb in general, but instead Mtb utilizes ESX-1 and PDIM to suppress the induction of a Th17 response that enhances protection against Mtb infection." I don't think the last sentence is necessarily true. I can imagine a scenario in which the induction of the Th17s is, in fact, due to the attenuation, and the Th17 induction still contributes to protection.

      We tested another attenuated M. tuberculosis strain with no known relationship with ESX-1 or PDIM, ΔMmpL4. This attenuated mutant fails to induce IL-17A–producing CD4 T cells to the same extent as observed in mice infected with ESX-1-deficient or PDIM-deficient strains, which is strong evidence that simple attenuation of virulence does not result in higher numbers of Th17 cells being elicited.

      (3) ESX-1, PDIM, and mmpl4 mutants all have similarly reduced CFUs in the lung, but what about the LN? The bacterial burden in the LN may be more important for regulating T-bet, IL-23, and Th17 differentiation, since the LN is where T cell priming occurs, than the CFU in the lung. Perhaps ESX-1 and PDIM mutants have reduced CFU in the LN, but mmpl4 does not. This difference in LN burdens may be the primary driver of Th17 priming, as high avidity interactions are thought to be an important driver of T-bet induction.

      We acknowledge that this is a formal possibility, however we maintain that the phenotype is specific to ESX and PDIM mutants, rather than MmpL4 mutants. Even if this phenotype arises from a tissue-specific attenuation of ESX/PDIM mutants, it remains a specific phenotype of these mutants, and not all attenuated mutants, albeit less directly. More importantly, the observation that these mutants induce higher levels of the Th17-polarizing cytokine IL-23 from infected cells ex vivo suggests that this is not an indirect phenomenon.

      (4) Do LN cDC1 and high levels of IL-12 p35 in mice infected with the mmpl4 mutant? Likewise, LN cDC2's express low levels of IL-12 p19 (akin to those infected with WT Mtb)? If these observations for ESX-1 and PDIM mutants are mechanistically linked to the increased numbers of Th17 cells, then you would expect mice infected with mmpl4 mutants to be more like those infected with WT Mtb than those infected with ESX-1 and PDIM mutants.

      Because ΔMmpL4 and complemented strains resulted in T cell profiles that were not different from the wild-type, we did not measure mediastinal lymph node dendritic cell expression of IL-12 p35 and IL-23 p19 in infections with these mutants.

      (5) ESX-1 and PDIM are very different virulence factors - a protein secretory pathway and cell wall lipid, respectively? Mechanistically, how would mutants in these pathways give very similar outcomes regarding Th17 cells unless it was simply as an aspect of their attenuation? Perhaps, mmpl4 mutants simply differ in some aspects of their attenuation, such as bacterial burdens in LNs, or their interaction with cDCs?

      We are not the first to link phenotypes of ESX-1 and PDIM. Both systems have both been shown to be important for M. tuberculosis permeabilization of the host cell phagosome after phagocytosis, and for suppression of type I IFN responses, among other responses. Thus, these seemingly different virulence factors clearly work together to support specific virulence traits during infection. The exact mechanism of how ESX-1 and PDIM interact is not completely understood and is an area for future investigation.

      Reviewer #2 (Public review):

      The following conclusions and interpretations should be revisited, rephrased, and re-evaluated:

      (1) The manuscript neglects to analyze T cell responses in the dLN, which is the critical site where these responses are initiated (only DC cytokine production is measured in the dLN). The differences in the lungs could reflect trafficking of T cells to the lungs, local lung T cell responses, or durability of the T cell responses in the lungs. The authors state in the last results section that "These results indicate that the ESX-1 and PDIM virulence factors impact naïve T cell differentiation at the draining mediastinal lymph node..." but T cell responses are never measured in the dLN.

      Due to the limited size of the mediastinal lymph node at 3 weeks post infection, we were unable to obtain enough cells for both myeloid cell analysis and T cell analysis, as we perform staining for these panels separately due to the decrease in viability of myeloid cells observed during T cell restimulation. In addition, because T cells in the lung are the population of cells most critical for mediating the outcome of infection, we believe analyzing the T cell response in the lymph nodes though interesting, is not crucial for this study. We have edited the manuscript to be clearer, as suggested by the reviewer.

      (2) Figure 2: The authors state that "Importantly, IFN-γ deficient mice did not exhibit elevated levels of IL-17A producing CD4 T cells demonstrating that IFN-γ production is not the mechanism by which Th1 T cells limit a Th17 response during Mtb infection", but the difference is significantly different and even more obvious in Panel B. In fact, if the Panel D y-axis was on a log scale, the Ifng-/- would likely look more like Tbet-/- than WT. Based on this data, it seems like IFNg is having an effect and should not be completely discounted. Does the deletion of Ifng affect the number of Tbet+ T cells?

      We agree that the IFN-γ<sup>-/-</sup> have only 5x more IL-17 producing CD4 T cells than WT mice while Tbet<sup>-/-</sup>mice exhibit a 25-fold increase compared to WT. We have added this information to the text, and now point out that IFN-γ production is not the sole mechanism by which Th1 T cells limit a Th17 response during Mtb infection.

      In addition, the deletion of Tbet results in an increased number of IFNg+IL-17+ double positive T cells (Figure 2B), in addition to a sizable IFNg single positive T cell population maintained in the Tbet-/- mice (10x the negative control of Ifng-/-). Is this why Tbet deletion is not as severe as Ifng deletion, because T cells are still making IFNg?

      It is possible that the residual IFN-γ produced by T-bet-deficient animals contributes to their relatively modest susceptibility to infection. However, our data show that deletion of IL-17 in this background renders T-bet–deficient mice nearly as susceptible as IFN-γ deficient mice, arguing that the remaining IFN-γ is not a major protective factor.

      Along these lines, the statement in the text that, "Tbet-/-Il17a-/- mice completely lacked both IFN-γ producing...." T cells is not supported by the data in Figure 2C. Tbet-/-Il17a-/- mice look to have more gamma-producing T cells than Tbet-/- mice (which is already 10x the negative control of Ifng-/- in panel 2B if one includes the gamma single positive and IFNg/IL-17 double positive).

      We have amended the language in the text to be more consistent with the data.

      (3) In the Results sections describing Figures 3, 4, and 5, the authors equate IL-17 production by T cells with TH17 responses and IFNg expression with TH1, but Tbet and RORgt expression in the T cells should be measured to make conclusions about TH1 and TH17. Or the authors can rephrase their findings to specifically state the observations as IFNg or IL-17 expressing CD4+ T cells.

      We believe that calling a CD4 T cell in the lung that is producing IFN-γ (and not IL-17) a Th1 cell is appropriate. Potentially confounding cells include those which also produce IL17, which we have ruled out, or T<sub>FH</sub> cells that may be common in lymph nodes but are not common in lungs at this time point and under these conditions.

      (4) Conceptually, do the authors think that ESX1/PDIM promotes TH1 responses and this blocks TH17 or are ESX1/PDIM blocking TH17 responses directly, allowing for increased TH1 responses? It would be helpful to clarify the model in this regard, describe how the data supports one model or the other, and then make sure the language is consistent throughout. Can these effects on T cell responses be tested and recapitulated in vitro using infected APC and T cell co-cultures?

      While it is possible that PDIM and ESAT-6 suppress Th17 through promotion of Th1 differentiation, we do not have data to support this model currently. However, we have added a comment making this point to the discussion.

      Reviewer #3 (Public review):

      Weaknesses:

      (1) The authors should acknowledge and reference key findings from the literature that have identified suppression of Th17 differentiation as an Mtb virulence mechanism, e.g., the role of the Hip1 protease and CD40 signaling (Madan-Lala JI 2014, Sia Plos Path 2017, Enriquez iScience 2022) and Khader JI 2005, showing the requirement of IL-23 for Th17 responses in vivo in a TB mouse model.

      We thank the reviewer for pointing these references out and have added them to the discussion section of the manuscript.

      (2) Addressing several questions related to the Tbet KO mouse experiments would strengthen the study. Do the Tbet KO mice have elevated IL-4/5/13 (which has been previously reported in non-TB studies) in addition to IL-17? The lack of Th17 cells in the IFNg KO compared to the Tbet KO may be due to a difference in timing, since only 3-week data are shown; earlier and later time points would provide better interpretation. The authors do not present any data on neutrophil infiltration in WT vs Tbet KO vs IFNg KO mice. Since IL-17 is known to be important for recruiting neutrophils to the lung, data on neutrophils are important for clarifying the mechanism for the CFU outcomes.

      We agree that it is surprising that, in the context of TB, Th17 responses are protective whereas excessive neutrophil recruitment is detrimental to the host. In IFN-γ–deficient mice, neutrophils are recruited and contribute to the increased susceptibility of this strain (PMID: 21967766). In separate work from our lab, we have shown that the phenotype of neutrophils recruited to the lungs during Mtb infection influences disease outcome (PMID: 40937719). It is possible that differences in the host environment and the timing of the response shape the effects of neutrophils on the host; these and the other questions raised by the reviewer will be the subject of future studies.

      (3) While IL-23 is important for sustaining IL-17 production, IL-6, TGF-b and/or IL-1β are necessary for Th17 polarization. What were the levels of these cytokines in DCs in the lung? (Figure 5). Additionally, Tbet-deficient DCs exhibit impaired activation of antigen-specific Th1 cells and have reduced IL-12 production. Given the data showing higher IL-17 levels in Tbet KO mice, the authors should provide information on the DC phenotype (IL-23, IL-6, etc.) in the Tbet KO experiments.

      While these are interesting points, investigating mechanisms of Tbet-dependent suppression of IL-17 is beyond the scope of this study.

      (4) The mechanism by which ESX-1/PDIM function to impact Th17 differentiation is not clear. While data showing a role for ESX-1 and PDIMs in inhibiting Th17 responses is interesting, there is no insight into the potential mechanism of action. Figure 3 showing reduction in IFNg+ CD4 T cells after infection with eccC1 and fadD28 mutants suggests that this outcome is due to a lower bacterial load relative to WT Mtb at the 3-week time point. Since IFNg is known to suppress IL-17, the higher levels of Th17 cells could be due to the reduction in IFNg due to the attenuated growth of the mutants. Additionally, what was the level of Type I IFNs elicited by these mutants?

      We included the MmpL4 knockout Mtb Erdman strain as a control to ensure that attenuation of mutants is not the cause of the increase in IL-17. We also showed that eliminating type I IFN signaling by deleting its receptor has minimal impact on Th17 differentiation, even in the context of a host that produces excess type I IFN. Therefore we do not believe that type I IFN elicited by these mutants is explanatory for the phenotype.

      (5) Since macrophages have been implicated in the reduced cytokines seen in the ESX-1 mutant, IL-23 and other cytokine data on lung macrophages would complement the DC data.

      Because dendritic cells are primarily responsible for priming CD4 T cell responses, we believe that this result in macrophages would not substantially alter our conclusions. That said, it was demonstrated previously that macrophages infected with ESX-1 mutants produce less IL-12p40, a subunit of IL-23.

      (6) Figure 5. There are many fewer DCs overall in the eccC1 and fadD28 mutant groups, which could account for the increased % IL-23p19 in DCs (5D). What were the levels of IL-23 in DC1s?

      The amount of IL-23 p19+ in type I conventional dendritic cells (cDC1s) was near zero as shown in supplementary figure 6A. cDC1s are known to not express IL-23 p19 in mice.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) What do the authors mean by the alternative secretion part of "ESX1 Type VII alternative secretion system" that they refer to?

      Bacterial alternative secretion systems facilitate the export of proteins from the bacterial cell independent of the canonical Sec-dependent secretion system required for export of most secreted bacterial proteins across the inner membrane. However to avoid confusion, we have removed the word alternative.

      (2) Not sure naïve fits in this sentence at the end of the introduction: "Furthermore, we observe a strong Th17 response during infection with ΔESX-1 or PDIM lacking Mtb in naïve mice....".

      We have removed the word Naïve.

      (3) Figure legend for 1A-C says analysis performed at 21 dpi, but the figure shows the time course.

      We have corrected this error.

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 1 should show the non-stimulated flow plot.

      We have added the unstimulated samples.

      (2) The % IL-17 in the flow plots is not consistent across Figures 1, 2, and 3. Not sure why the scales for the Y-axis for IL-17 differ so much between Figures 1 and 2/3. IS there a technical issue with compensation?

      We did not experience any difficulties with compensations. These experiments were done over several years of work. For every experiment, new single-color controls were used and gating was done with the FMO gating strategy. Minor variation such as we see here is not surprising.

      (3) Discuss Yeh et al J Neuroimmunol 2014- show that IFNγ inhibits Th17 differentiation and function via Tbet-dependent and Tbet-independent mechanisms.

      We have added this reference to the manuscript.

    1. eLife Assessment

      This useful study reports findings that support the use of the Open Field Test in Drosophila as a model to study "emotion-like states", which are behavioral responses to several stressful or aversive treatments, and resilience upon their subsequent removal. Behavioral data, by employing established stress-causing treatments and genetic manipulations, are solid. The main advance of this work over previous Drosophila work using a similar experimental setup is systematic analysis of various stressors.

    2. Reviewer #1 (Public review):

      Summary:

      Animal behavior is continuously influenced by the internal state moment by moment, including emotion primitives as the authors pointed out. Although emotion is a more human-related state, evolutional conservation is undeniable, which can be inferred by the behavioral manifestation. To further elaborate the neuronal mechanisms of emotion primitives, the simplest behavioral parameter related to emotional primitives should be well characterized. In this study, the authors described in detail of wall-following behavior (WAFO) and the total walking distance (TOWA) using flies after subjecting them to various conditions or flies being genetically manipulated according to the previous reports that could affect emotion primitives. Overall, the study is well designed and structured. In addition, the discussion on emotion primitives will be of value to the field.

      Strengths:

      The strength of this study is its use of a simple behavioral parameter, TOWA, and also a simple design of behavior, WAFO. The importance of the behavioral assay is reproducibility and comparability. In fact, the author demonstrated a summary of comparisons where different treatments result in scalable behavioral changes in WAFO and TOWA.

      Conceptual concerns:

      My suggestion to strengthen the authors' conclusion that "TOWA can be interpreted as a behavioral proxy for exogenously induced arousal" was to show that an increase in TOWA after stress exposure can be observed in a small (1-cm) arena that acts as an exogenous arousing stimulus, but not in a larger arena (>6.6 cm) where such arousing effects are absent. This comparison would demonstrate that basal locomotor activity measured in larger arenas is not altered by stress, whereas the additional component observed in smaller arenas reflects stress-induced internal state. Therefore, the authors would be able to distinguish clearly the effects of stressors or experiences on either simple locomotion or an emotion-like internal state. Then the future works can follow this protocol using smaller and larger arena to assess emotion-like internal state.

      I appreciate the significant authors' efforts to monitor TOWA using arenas with different diameters up to 6 cm. However, the conclusion was unfortunately the same as that obtained using the 1-cm arena. As the authors commented, flies do not show persistent and quantifiable wall-following in arenas larger than 5.8 cm, which limits further examination of this question. I personally agree with the authors' interpretation, but I hope that the authors obtain more definitive experimental contrasts to support this claim in the future study.

    3. Reviewer #2 (Public review):

      Summary:

      In terms of data, the revised manuscript is by and large the same, though the authors added new experiments examining the effects of the Open Field Test (OFT) arena size (Fig. 1 Supplement 1) and sex and mating status (Fig. 6). The authors have also provided textual revisions, partially addressing my previous major criticism about novelty over the work of Mohammad et al., 2016, Curr Biol. They argue that the main advance is the systematic inclusion of Total Walking (TOWA) data (e.g. Introduction, page 6 top in the tracked changes document). While I am still not convinced that the findings represent a huge leap forward over that previous work, the authors' systematic analysis is very nice and may prove of use to those seeking to develop Drosophila as a model for studying emotion primitives.

      Strengths:

      The main strength of the paper is the rigorous use of several stressful or aversive treatments and their subsequent removal to show that WAFO is a robust proxy for stress-like emotional primitives across multiple stimuli. The pharmacological, molecular, and neuronal activity manipulations, although more limited in scope, lend further credence to the authors' central claim.

      Weaknesses:

      The authors have addressed some of my previous points with textual revisions and in their rebuttal. As stated above, the conceptual advance over Mohammad et al. remains in my opinion limited, but I appreciate that this point is now clearly discussed in the manuscript.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      The strength of this study is its use of a simple behavioral parameter, TOWA, and also a simple design of behavior, WAFO. The importance of the behavioral assay is reproducibility and comparability. In fact, the author demonstrated a summary of comparisons where different treatments result in scalable behavioral changes in WAFO and TOWA.

      We appreciate this assessment and fully agree that the simplicity of the assay and the demonstration of its scalability and reproducibility are a strength.

      Weaknesses:

      The weakness of the study is the lack of further experiments to support their assumption related to TOWA. The authors suggested that TOWA can be interpreted as a behavioral proxy for exogenously induced arousal. However, it could be interpreted as higher activity, although the authors argued that the circadian clock increasing locomotor activity around ZT0 and ZT12 does not affect TOWA, and therefore TOWA is not related to the locomotor activity per se. As the author cited, flies lose locomotor activity in the circular arena of 6.6 cm in diameter, whereas they continuously move during a 1-h recording in the authors' arena of 1 cm in diameter.

      I would agree that the arena of 1 cm in diameter, but not 6.6 cm in diameter, serves as an exogenous stimulus inducing arousal, and TOWA is manifested by arousal. However, TOWA would also be affected by other behavioral parameters, including the activity, motivation for exploration, or perception of the space. Therefore, it could be reasonable to re-examine some of the flies tested in this study in the circular arena of 6.6 cm in diameter. If arousal is biased by the components presented in Figure 6 and TOWA can assess mainly exogenously induced arousal, the treatment altering TOWA in the arena of 1 cm in diameter would not affect their behavior in the arena of 6.6 cm in diameter. My concern is that Figure 6 may demonstrate too simplistic a diagram to interpret the results. I would suggest adding the experiments using the arena of 6.6 cm diameter or softening the argument.

      We are grateful that you prompted us to investigate the relation between TOWA and arousal and different arena diameters in more depth. Based on your comments, we compared naïve and stressed behaviour between arena diameters of 1, 2.2 and 5.8 cm. The sizes were chosen as to optimally comply with the camera field of view in our setup. Naïve flies showed stable locomotor activity throughout the 60 min of recordings in the different arenas (new Figure 1 – figure supplement 1 A-B’’). Moreover, no significant difference in TOWA over the first 10 min was found between the different arena sizes (new Figure 1 – figure supplement 1 C). This suggests to us that arenas with a diameter smaller than 6.0 cm (and not only with a 1 cm diameter) induce some form of activity that resembles stimulated activity as defined by Meehan and Wilson (1987). A mechanical shock (shake) resulted in significantly increased WAFO in all arena sizes (new Figure 1 – figure supplement 1 D-D’’). We did not test the effects in larger arenas > 5.8 cm, as we found that flies do not longer show persistent and quantifiable wall following. We started to see that also in few flies in the 5.8 cm arena – these few flies were excluded from our analysis presented in new Figure 1 – figure supplement 1. Unlike the naïve response, the stress-induced response in WAFO and TOWA appears to be transient (new Figure 1 – figure supplement 2), which is in line with the definition of emotions as a transient state.

      Reviewer #2 (Public review):

      Summary:

      Strengths:

      The main strength of the paper is the rigorous use of several stressful or aversive treatments and their subsequent removal to show that WAFO is a robust proxy for stresslike emotional primitives across multiple stimuli. The pharmacological, molecular, and neuronal activity manipulations, although more limited in scope, lend further credence to the authors' central claim.

      We are glad about this assessment and share your opinion.

      Weaknesses:

      The conceptual advance of this research is unclear, as previous work (Mohammad et al., 2016, Curr Biol.) carried out similar treatments and manipulations and reached largely similar conclusions.

      Thank you very much for bringing this up. We rewrote respective parts of the introduction (second last paragraph) and discussion to more clearly outline the advances over the previous work by Mohammad et al. 2016. While our study builds upon Mohammad et al. 2016, the conceptual advance and novelty is that we constitute and treat TOWA as a second and independent dimension equal to WAFO in the OFT. Mohammad et al. had measured locomotor activity (reported as average speed in their paper (total distance walked/time of recording), but primarily to test the dependency of WAFO on locomotor activity. They found that WAFO metrics were poorly correlated with average walking speed, showing a significant degree of independence of both measures – a finding that our results confirm. However, unlike us, they did not consider average speed/TOWA further for their analysis, possibly because they focused on anxiety-like behaviour while our study looked broader on emotion-like behaviour in general. We further used a round (not square) arena to exclude “cornering” in order to reduce the complexity of the assay, which may explain differences of observed speed/TOWA between our studies.

      Moreover, while WAFO is a good proxy for 'stress', I am not convinced that TOWA necessarily represents an emotional state in all cases. Indeed, as the authors themselves acknowledge, changes in total walking may be associated with other factors, such as starvation-induced hyperactivity, physical exhaustion after sleep deprivation, increased sex drive after mating, alcohol sedation, etc.

      Your comment raises a question in comparative research on emotions which is very difficult if not impossible to conclusively answer. At first sight, the most conservative stance seems to be to completely disregard the idea of emotions and affective experiences in animals. This, however, would mean that we cannot use animal models to study the basics of emotions (= emotion primitives) and would ignore that by all likelihood emotions are a product of evolution and hence should exist at least in more basic forms in animals. Obviously, we have no means to ask flies or any other animal whether they connect “hunger” or “mating” to a feeling or an emotional state (which must not be conscious) but can only observe the behaviour. We here adopt the often-cited “Pankseppian” view (based on the book of Jaak Panksepp: “Affective Neuroscience”) and firmly believe that – in order to fully understand how the brain drives behaviour- we also need to take affective states into account that bias behaviour towards adaptive responses.

      In short, we are unfortunately unable to give a clear and definite answer to your comment whether TOWA represents an emotional state in all cases. Perhaps you are right. We believe, however, that a “Pankseppian” view is adequate, and we may ask what evidence exists that shows that starvation-induced hyperactivity or post-mating is not associated with an affective emotion-like state in the fly or any other animal.

      Another unclear point is the interpretation of some unexpected results, such as the finding that both serotonin transporter overexpression and its knockdown give the same phenotype.

      Thank you very much for this comment, which we also received by reviewer #1. As suggested by the other reviewer, a reasonable mechanism that may underly the similarity of effect after knockdown or overexpression of SerT may be a differential effect on the different serotonin receptor subtypes expressed in the brain. The possibility that the concentration-dependent effect of a biogenic amine follows a U-shape is further reasonable and has been demonstrated for dopamine. For example, the relationship between cognitive performance or working memory and dopamine levels in primates follows an inverted U-shape (see e.g. Cools and D’Esposito 2011 10.1016/j.biopsych.2011.03.028, Desimone 1995 10.1038/376549a0). Also in Drosophila, both reduced and increased dopamine levels lead to increased male-to-male courtship behaviour (Liu et al. 2008 10.1523/JNEUROSCI.5290-07.2008, Liu et al. 2009 10.1371/journal.pone.0004574). While we are unaware of similar examples for serotonin, we note that serotonin levels must be kept at optimum level during development – both higher and lower than optimum levels result in behavioural impairments in adults (see e.g., Shah et al. 2018 10.3389/fnbeh.2018.00114). We have now extended the discussion accordingly.

      Finally, there are some issues with the use of the OFT in rodent research (e.g., inconsistent effects of anxiolytic drugs; see Rosso et al., 2022, Neurosci Biobehav Rev., for a meta-analysis). These should be explained to place the Drosophila findings in their appropriate context.

      Thank you very much for bringing this systematic review to our attention which assessed the usefulness of various behavioural tests including the OFT to study the effect of anxiolytic drugs in rodents. Overall, the review casts “serious doubt on both construct and predictive validity” of behavioural tests for anxiolytics. While diazepam (the only drug used in our study) was the drug with the most consistent effects across the analysed behavioural assays, only 59% of the OFTs revealed significant effects. We were already aware of earlier findings in the same direction (Prut and Belzung 2003 10.1016/s0014-2999(03)01272-x), but as we only used one drug did not include a discussion in the manuscript. Unfortunately, the number of studies employing the OFT in flies is very small and does not yet allow for a similar comparison. We now changed the respective sentences in the discussion:

      “In rodents, diazepam mostly but not consistently leads to an anxiolytic response in the OFT behaviour which questions the usefulness of the OFT for testing anxiolytic drugs (see (Prut and Belzung 2003; Rosso et al. 2022)).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Overexpression of SerT suppressed increased WAFO after electric shocks. However, knockdown of SerT did similar. It is worth reporting this finding, but possible mechanisms could be mentioned in the manuscript. For example, given that flies carry five serotonin receptor genes, upregulation of overall serotonin level may affect the specific serotonin receptor, but downregulation of it may affect other receptors, leading to unexpected outcomes. Exploring any other possibilities could be better to add to guide future research.

      Thank you very much for this comment and the suggestion of a reasonable mechanism that may underly the similarity of effect after knockdown or overexpression of SerT. The possibility that the concentration-dependent effect of a biogenic amine follows a U-shape is further reasonable and has been demonstrated for dopamine. For example, the relationship between cognitive performance or working memory and dopamine levels in primates follows an inverted U-shape (see e.g. Cools and D’Esposito 2011 10.1016/j.biopsych.2011.03.028, Desimone 1995 10.1038/376549a0). Also in Drosophila, both reduced and increased dopamine levels lead to increased male-to-male courtship behaviour (Liu et al. 2008 10.1523/JNEUROSCI.5290-07.2008, Liu et al. 2009 10.1371/journal.pone.0004574). While we are unaware of similar examples for serotonin, we note that serotonin levels must be kept at optimum level during development – higher or lower levels result in behavioural impairments in adults (see e.g. Shah et al. 2018 10.3389/fnbeh.2018.00114). We have now extended the discussion accordingly.

      Reviewer #2 (Recommendations for the authors):

      (1) The advance over Mohammad et al., 2016, Curr Biol. must be clearly and emphatically articulated in the Introduction and/or Discussion. It is otherwise impossible to appreciate what is conceptually novel about this work.

      We have now rewritten parts of the two last paragraphs in the introduction to make the advances clearer. As outlined above, we conceptually advanced the analysis of OFT behaviour by integrating TOWA as a second and independent dimension in our analysis. During the revision process, we have spent great effort to better characterise the nature of the locomotor activity encountered in the OFT (Figure 1 – supplementary figures 1 and 2). Also, this is an advancement over previous studies, including Mohammad et al. 2016. We further included new results on the general effect of neuropeptides (silver mutants, impaired in neuropeptide processing).

      (2) The most important metric for a stress-like emotional primitive is WAFO. Therefore, figures should be revised in a way that highlights WAFO differences. Figure 4 is a good example of this. In contrast, in Figures 1-3, the WAFO box-and-whisker plots are very small and obscured under the raw WAFO-TOWA plots, which are difficult to see (especially given the light blue background of all plots) and redundant. I strongly recommend just showing the WAFO box-and-whisker plots for the sake of visibility, clarity, and brevity.

      Although we understand your reasoning, we would like to stick with the old figures as we consider TOWA as important to characterise the OFT response as WAFO (see our comments above regarding the conceptual advances of our study over Mohammad et al. 2016). It is true that Figures 1-3 are small, but at least in our print-out well legible. Further, it appears that in the current version of the eLife system the resolution is downsampled. In addition, we anticipate that in the version of record figures can be enlarged online as in other eLife articles.

      (3) Given the criticisms against the OFT in rodent research, one of which is inconsistent effects of anxiolytic drugs (Rosso et al., 2022, Neurosci Biobehav Rev.), it may be useful to expand pharmacological treatments beyond diazepam.

      As our focus is not on the testing of anxiolytic drugs and since it was already very difficult to be granted access to diazepam (we are not at a medical institution), we refrained from testing further drugs. Moreover, as rightfully mentioned by you, the OFT may not be the best test for the efficacy of anxiolytic drugs. On the other hand, diazepam was the most consistent anxiolytic in the OFT in rodents (see Rosso et al. 2022).

      (4) The authors should show results of the effects of at least some stressors/punishments on WAFO/TOWA of female flies to understand if observed effects are sex-specific.

      Thank you very much for bringing this topic to our attention. To test whether the effects are sex-specific, we now performed several new experiments. First, we compared the naïve OFT response of mated and unmated males and females (see new Figure 6). This revealed that without prior stress treatment, the WAFO response is independent of sex and mating status. In contrast, the naïve TOWA response turned up to be sex- and mating state-specific (see new Figure 6). To test whether the mated females show a different stress-induced OFT response to males, we applied mechanical stress (shake) that we had also used to assess the effects of arena diameter (Figure 1 – supplementary Fig. 1 D-D’’). After a first round of shaking, females showed increased TOWA, but WAFO was unaffected. A second round of shaking, however, led to a significant increase in WAFO and TOWA. This suggests that the OFT response is qualitatively similar between the sexes and mating status, yet the threshold for elicited responses differs between males and females. We now added a respective paragraph to the main text in the results section plus a new figure (Fig. 6).

    1. eLife Assessment

      This important study addresses a timely issue at the intersection of mitochondrial and telomere biology by focusing on the relationship between naturally occurring variants in the mitochondrial genome and telomere length. This work thereby provides a conceptual and experimental framework for investigating communication between mitochondria and telomeres. Using an innovative transmitochondrial cybrid approach, the authors provide evidence that mitochondrial DNA variants influence telomere maintenance through effects on mitochondrial function, reactive oxygen species, and NAD⁺-dependent repair processes. The evidence supporting the central conclusion that mitochondrial genotype influences telomere-associated phenotypes is convincing and is strengthened by the use of complementary functional and rescue experiments. However, some of the mechanistic interpretations and broader conclusions regarding telomere length inheritance in humans would benefit from additional donors and longitudinal analyses following cybrid generation, or more cautious framing.

    2. Reviewer #1 (Public review):

      Summary:

      This is an interesting study that addresses whether mitochondrial DNA (mitoDNA) variants impact telomere length (TL), which may be relevant to potential maternal inheritance of TL in offspring. The study addresses this question using a cybrid model approach in which mitochondria from donor platelets from 7 individuals that vary in TL and differ in mitoDNA variants are introduced into 143B cells that lack mitochondria. MitoDNA variants that exhibited reduced complex I activity showed telomere shortening in cybrids and increased telomere dysfunction. Interestingly, these phenotypes could be reduced with NAC antioxidant and NAD+ supplementation, suggesting that ROS and oxidative DNA damage at telomeres contributed to the telomere shortening. They further showed that cybrids with lower levels of ROS correlated with longer TL in the lymphocytes of the mitochondrial donors.

      Strengths:

      This study provides compelling evidence that mtDNA variants influence TL through a mechanism involving mitochondrial-derived ROS, potentially causing telomeric oxidative damage. The data are robust, and the manuscript is well written. However, the study could be strengthened by addressing the following questions and minor weaknesses below.

      Weaknesses:

      (1) Introduction. Line 81, the relationship between TL and the risk of lymphoid and myeloid leukemia is not straightforward. POT1 variants associated with long TL increase the risk for lymphoid and myeloproliferative neoplasms (see PMID: 41564438 for example).

      (2) Figure 1. Since sex also influences TL, it would be good to know the sex of the selected individuals or explain why this is not necessary.

      (3) Please include a description of the 143B cells that were used for cybrid formation in the Results section when introducing the cybrids.

      (4) Lines 155-156. The authors note that cybrids from donors 1 and 2 show "pronounced" telomere damage. This result indicates an increase in 53BP1-positive telomeres, which could be indicative of telomere dysfunction or damage. Quantification of the increased chromosome end fusions for cybrids 1 and 2 would strengthen the result. Do the increased fusions correlate with an increase in telomere signal-free ends? These should be apparent in the telomere FISH images of metaphase chromosomes.

      (5) Lines 168-169. What is the evidence that the "in vitro metabolic shift" causes acute oxidative stress?

      (6) Why did the elevated ROS in cybrid #3 (Figure 4C) not translate to shorter telomeres in the cybrid (Figure 2A)? Perhaps there is a difference between factors that determine TL in the cybrid vs the donor's lymphocytes? In Figure 4B, it appears that the statistical comparisons for mitochondrial superoxide are all relative to Cyb3. If so, why are the comparisons not with the parental 143B rho0 cell line? Please clarify.

      (7) Given the heterogeneity in TL and mtDNA variants in the human population, the conclusions could be further strengthened by increasing the number of donors and cybrids analyzed. However, there are admittedly practical factors. Overall, these findings are compelling and provide a solid foundation for expanding this analysis in the future. This is more of a comment than a weakness.

    3. Reviewer #2 (Public review):

      Summary:

      The authors aim to determine whether mitochondrial genotype influences telomere length. By generating cybrids harboring different mitochondrial backgrounds, the authors seek to establish a mechanistic link between mitochondrial status and telomere biology.

      Strengths:

      A major strength of the study is the use of cybrid technology, which provides a great approach to investigate the role of mitochondrial DNA independently of the nuclear genome. The authors also employ multiple complementary assays to assess telomere-related phenotypes associated with mitochondrial dysfunction. Together, these experiments generate an interesting dataset that will be of value to researchers interested in the intersection between mitochondrial biology, genome stability, aging, and development. These results also build on previous work supporting roles for ROS/mitochondria in driving telomere shortening.

      Weaknesses:

      The data support the conclusion that mitochondrial background is associated with differences in telomere length and telomere-related phenotypes. However, some of the mechanistic interpretations would benefit from additional evidence. In particular, the manuscript discusses mitochondrial influences on telomere shortening, yet telomere length in some experiments is assessed at a single time point. Consequently, the current data do not directly address the rate of telomere attrition. Differences observed between cybrid lines could potentially arise from events occurring during cybrid formation, clonal selection, or subsequent cell expansion. Longitudinal analyses across multiple passages, ideally beginning immediately after cybrid generation and controlling for population doublings, would help establish whether mitochondrial function directly affects telomere shortening dynamics. Some experimental results would also benefit from additional quantification, clarification, and some biological replicates are missing.

      Overall, this study provides interesting evidence linking mitochondrial background to telomere biology. The cybrid models represent a useful resource for the field, and the work raises important questions regarding mitochondria-telomere communication.

    4. Reviewer #3 (Public review):

      Strengths:

      Mahieu and colleagues address an interesting and underexplored question: whether non-pathogenic variation in the mitochondrial genome contributes to the inter-individual variability of human telomere length (TL). Using a Belgian Flow-FISH reference cohort (n=491) to identify donors at TL extremes, they generate transmitochondrial cybrids from platelets of seven donors of distinct mtDNA subhaplogroups and characterize the resulting cells with a broad and well-executed toolkit (TRF, TeSLA, ddTRAP, EPR-based mitoROS, Seahorse with permeabilized-cell ETC dissection, LC-MS metabolomics, telomeric PAR-FISH). The most compelling finding is that cybrids derived from donors with low complex I (CI) activity undergo rapid telomere shortening during the glycolysis-to-OXPHOS transition of cybrid formation, and that this is largely prevented by co-treatment with NAC and the NAD⁺ precursor nicotinamide riboside, supporting a model in which CI sustains the NAD⁺ pool required for PARP1-mediated repair of oxidative damage at telomeres. The authors further report an inverse correlation between donor lymphocyte TL and mitoROS in the corresponding cybrids, and provide preliminary evidence that the K1a-defining ATP6 A177T variant (m.G9055>A) may be enriched in long-telomere individuals.

      Weaknesses:

      (1) Statistical support and donor sampling for the central in vivo correlation (Figure 4C).

      The inverse correlation between donor lymphocyte TL and cybrid mitoROS (R²=0.794, p=0.007) is the principal in vivo claim of the paper, but it is built on seven donors deliberately selected from the extremes of the Flow-FISH distribution. Sampling at the tails of the outcome variable can substantially inflate apparent correlation strength and significance. I would encourage the authors to (i) explicitly state this sampling structure where the correlation is introduced, (ii) report a leave-one-out sensitivity analysis to confirm the relationship is not driven by one or two donors (Cyb3 and Cyb6 appear to anchor the line), and (iii) where feasible, extend the analysis to additional donors with intermediate TL to test whether the relationship holds across the full distribution. Even a modest expansion (e.g., 4 to 5 additional donors at P25 to P75) would substantially strengthen this central claim.

      (2) Reconciling the cybrid CI / TL relationship (Fig 3B) with the absence of a CI / TL relationship in donor lymphocytes (Figure 4A).

      Figure 3B shows a strong correlation between CI activity and TL in cybrids (R²=0.87), while Figure 4A shows no correlation between donor CI activity (measured in the same cybrids) and donor lymphocyte TL. The authors acknowledge this, but the manuscript subsequently builds toward a CI-centric model of in vivo TL regulation, which seems to outrun the data. The most internally consistent interpretation is that the cybrid CI phenotype reports a sensitized in vitro response to the acute oxidative stress of the metabolic shift, rather than a steady-state determinant of leukocyte TL. I would suggest reframing the abstract, significance statement, and Discussion to make this distinction clearer. The in vitro CI / NAD⁺ / PARP1 axis is a strong finding on its own, while the in vivo role of CI activity (as opposed to ROS more broadly) is not yet established here. Donor #1's profile (very long lymphocyte TL, low CI activity, severe shortening in cybrids, no telomere inheritance in offspring) is informative in this regard and could be discussed more directly as a case that helps delineate where the cybrid model does and does not recapitulate in vivo biology.

      (3) The K1a / ATP6 A177T inheritance claim.

      The proposal that K1a (and specifically ATP6 A177T) contributes to maternal inheritance of long telomeres is intriguing but currently rests on three pedigrees (one of which, donor #1, does not support the hypothesis) and a chi-square test that does not reach significance (p=0.153, Figure 4F). The supporting evidence is also limited by the fact that platelet-mediated mitochondrial transfer delivers donor mitochondrial proteins, lipids, and residual mtRNA in addition to mtDNA, making it difficult to attribute the cybrid phenotype of donor #6 specifically to the ATP6 A177T variant. I would recommend either: (a) extending the genotyping screen to additional unrelated donors and, if feasible, confirming the effect of ATP6 A177T through an isogenic approach (e.g., mtDNA base editing in a clean background), or (b) softening the relevant statements to "suggestive trend warranting larger studies," and presenting the K1a observation as hypothesis-generating rather than supportive. The Ashkenazi-centenarian connection raised in the Discussion is an excellent direction for follow-up and could be framed accordingly.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an interesting study that addresses whether mitochondrial DNA (mitoDNA) variants impact telomere length (TL), which may be relevant to potential maternal inheritance of TL in offspring. The study addresses this question using a cybrid model approach in which mitochondria from donor platelets from 7 individuals that vary in TL and differ in mitoDNA variants are introduced into 143B cells that lack mitochondria. MitoDNA variants that exhibited reduced complex I activity showed telomere shortening in cybrids and increased telomere dysfunction. Interestingly, these phenotypes could be reduced with NAC antioxidant and NAD+ supplementation, suggesting that ROS and oxidative DNA damage at telomeres contributed to the telomere shortening. They further showed that cybrids with lower levels of ROS correlated with longer TL in the lymphocytes of the mitochondrial donors.

      Strengths:

      This study provides compelling evidence that mtDNA variants influence TL through a mechanism involving mitochondrial-derived ROS, potentially causing telomeric oxidative damage. The data are robust, and the manuscript is well written. However, the study could be strengthened by addressing the following questions and minor weaknesses below.

      We thank the reviewer for this very positive evaluation of our work.

      Weaknesses:

      (1) Introduction. Line 81, the relationship between TL and the risk of lymphoid and myeloid leukemia is not straightforward. POT1 variants associated with long TL increase the risk for lymphoid and myeloproliferative neoplasms (see PMID: 41564438 for example).

      We appreciate the reviewer raising the important link between pathogenic POT1 variants and lymphoid malignancies driven by elongated telomeres. We would like to clarify, however, that our introduction focused not on the pathological telomere attrition characteristic of telomere biology disorders, but rather on the natural, non-pathological variations observed in individuals with baseline telomere lengths on the shorter end of the spectrum.

      Nevertheless, we will include this observation regarding patients with long telomeres in the introduction to underscore the complex relationship between telomere length and tumorigenesis.

      The two following publications by the Armanios lab will be added:

      DeBoy EA et al. Familial clonal hematopoiesis in a long telomere syndrome. 2023. N Engl J Med 388, 2422-2433.

      Davidson-Swinton HR et al. Lymphoid malignancy and clonality in the POT1-mediated long telomere syndrome. 2026. Blood 147, 2226-2237.

      (2) Figure 1. Since sex also influences TL, it would be good to know the sex of the selected individuals or explain why this is not necessary.

      Because this study investigated the potential maternal inheritance of TL via the mitochondrial genome, our cohort consisted predominantly of female donors (6/7). Donor #5, the husband of Donor #7, was the only male included. We will update Figure 1D to include the sex of each donor and clarify this rationale in the text.

      Notably, our analysis revealed minimal influence of sex on TL, which cannot account for the observed differences between the extreme groups.

      (3) Please include a description of the 143B cells that were used for cybrid formation in the Results section when introducing the cybrids.

      We will do so.

      (4) Lines 155-156. The authors note that cybrids from donors 1 and 2 show "pronounced" telomere damage. This result indicates an increase in 53BP1-positive telomeres, which could be indicative of telomere dysfunction or damage. Quantification of the increased chromosome end fusions for cybrids 1 and 2 would strengthen the result.

      We thank the reviewer for this suggestion. Following their advice, we used our metaphase spread FISH analyses to quantify chromosome end fusions in the parental 143B Rho0, Cybrid 1 and Cybrid 6 cells. However, because Cybrid 2 metaphase spreads were of insufficient quality for adequate chromosome analysis, we will restrict our telomere fusion comments exclusively to Cybrid 1 and include the quantifications in our revised manuscript.

      Author response image 1.

      The number of fusions and total chromosomes analyzed are indicated on each bar.

      Do the increased fusions correlate with an increase in telomere signal-free ends? These should be apparent in the telomere FISH images of metaphase chromosomes.

      While this is a strong argument, the telomeres within these cybrid models are critically short. Consequently, the FISH signal intensity falls below the threshold required for reliable quantification of telomere-free ends.

      (5) Lines 168-169. What is the evidence that the "in vitro metabolic shift" causes acute oxidative stress?

      The reviewer is correct that we have not formally demonstrated this in our experimental system. Instead, our assumption was based on the fact that the initial phase of Rho0 cell repopulation involves a temporary ROS burst, partly driven by incompletely assembled ETC supercomplexes that are known to elevate ROS levels (Maranzana et al, 2013). This is supported by our experimental observation that the NAC antioxidant, combined with NR, potently inhibits telomere shortening in cybrids with low CI activity. We propose to add this explanation in the revised manuscript.

      (6) Why did the elevated ROS in cybrid #3 (Figure 4C) not translate to shorter telomeres in the cybrid (Figure 2A)? Perhaps there is a difference between factors that determine TL in the cybrid vs the donor's lymphocytes?

      We cannot fully explain the discrepancy, but we indeed suspect in vivo oxidative stress differs significantly from cell cultures (21% O<sub>2</sub>). Additionally, early cybrid replenishment involves an unknown telomere elongation step that may offset mitochondrial ROS-induced shortening.

      To further investigate this question, we measured telomere length across varying population doublings (PDs) and observed the following shortening after 66-67 PDs:

      - Cybrid 3: about 1.6 kb reduction

      - Cybrid 6: about 1.2 kb reduction

      - Cybrid 7: about 0.8 kb reduction

      These results suggest that the rate of telomere shortening in culture may be higher in Cybrid 3 cells, possibly due to increased ROS levels.

      We propose to include (as Supplementary figure) and discuss these data in the revised manuscript.

      Author response image 2.

      In Figure 4B, it appears that the statistical comparisons for mitochondrial superoxide are all relative to Cyb3. If so, why are the comparisons not with the parental 143B rho0 cell line? Please clarify.

      We excluded Rho0 cells from our superoxide measurements because these cells were grown in a different culture medium. Instead, we focused on comparing mitochondrial ROS across cybrids to accurately correlate these values with the telomere length of the corresponding donors’ lymphocytes. Consequently, Rho0 cell measurements would not have contributed to this correlation analysis.

      We propose to discuss this in the revised manuscript.

      (7) Given the heterogeneity in TL and mtDNA variants in the human population, the conclusions could be further strengthened by increasing the number of donors and cybrids analyzed. However, there are admittedly practical factors. Overall, these findings are compelling and provide a solid foundation for expanding this analysis in the future. This is more of a comment than a weakness.

      We thank the reviewer for this constructive feedback and entirely agree that including more donors would have added depth to our findings. While we acknowledge this limitation, pursuing this further is currently impossible without acquiring new ethical approvals and establishing fresh collaborations with clinicians.

      Reviewer #2 (Public review):

      Summary:

      The authors aim to determine whether mitochondrial genotype influences telomere length. By generating cybrids harboring different mitochondrial backgrounds, the authors seek to establish a mechanistic link between mitochondrial status and telomere biology.

      Strengths:

      A major strength of the study is the use of cybrid technology, which provides a great approach to investigate the role of mitochondrial DNA independently of the nuclear genome. The authors also employ multiple complementary assays to assess telomere-related phenotypes associated with mitochondrial dysfunction. Together, these experiments generate an interesting dataset that will be of value to researchers interested in the intersection between mitochondrial biology, genome stability, aging, and development. These results also build on previous work supporting roles for ROS/mitochondria in driving telomere shortening.

      We thank the reviewer for this very positive evaluation of our work.

      Weaknesses:

      The data support the conclusion that mitochondrial background is associated with differences in telomere length and telomere-related phenotypes. However, some of the mechanistic interpretations would benefit from additional evidence. In particular, the manuscript discusses mitochondrial influences on telomere shortening, yet telomere length in some experiments is assessed at a single time point. Consequently, the current data do not directly address the rate of telomere attrition. Differences observed between cybrid lines could potentially arise from events occurring during cybrid formation, clonal selection, or subsequent cell expansion. Longitudinal analyses across multiple passages, ideally beginning immediately after cybrid generation and controlling for population doublings, would help establish whether mitochondrial function directly affects telomere shortening dynamics. Some experimental results would also benefit from additional quantification, clarification, and some biological replicates are missing.

      We limited our TL measurements to the earliest viable time point after cybrid formation to avoid the confounding effects of cellular adaptation in culture. For instance, the early telomere shortening observed in Cybrid 1 and Cybrid 2 was later alleviated—likely due to the upregulation of the NAD+ salvage pathway genes NAMPT and NAPRT1.

      To accurately capture the effects specific to cybrid formation, we isolated 4 independent clones per donor, all of which showed highly consistent TL values, as shown in Figure 2A and S3D.

      While we acknowledge the reviewer's point about multi-passage longitudinal analyses, we did, in fact, measure telomere length across varying population doublings (PDs) in Cybrid 3 (high mito ROS levels), 6 and 7 (low mito ROS levels) and observed the following shortening after 66-67 PDs:

      - Cybrid 3: about 1.6 kb reduction

      - Cybrid 6: about 1.2 kb reduction

      - Cybrid 7: about 0.8 kb reduction

      These results suggest that the rate of telomere shortening in culture may be higher in Cybrid 3 cells, possibly due to increased ROS levels.

      We propose to include a Supplementary figure and discuss these data in the revised manuscript.

      We further propose to carefully edit the manuscript so as to clarify the text and, whenever possible, add quantifications. Among others, as suggested by Reviewer #1, we will add the quantification of chromosome end fusions in the parental 143B Rho0, Cybrid 1 and Cybrid 6 cells in our revised manuscript (See Author response image 1).

      Overall, this study provides interesting evidence linking mitochondrial background to telomere biology. The cybrid models represent a useful resource for the field, and the work raises important questions regarding mitochondria-telomere communication.

      Reviewer #3 (Public review):

      Strengths:

      Mahieu and colleagues address an interesting and underexplored question: whether non-pathogenic variation in the mitochondrial genome contributes to the inter-individual variability of human telomere length (TL). Using a Belgian Flow-FISH reference cohort (n=491) to identify donors at TL extremes, they generate transmitochondrial cybrids from platelets of seven donors of distinct mtDNA subhaplogroups and characterize the resulting cells with a broad and well-executed toolkit (TRF, TeSLA, ddTRAP, EPR-based mitoROS, Seahorse with permeabilized-cell ETC dissection, LC-MS metabolomics, telomeric PAR-FISH). The most compelling finding is that cybrids derived from donors with low complex I (CI) activity undergo rapid telomere shortening during the glycolysis-to-OXPHOS transition of cybrid formation, and that this is largely prevented by co-treatment with NAC and the NAD⁺ precursor nicotinamide riboside, supporting a model in which CI sustains the NAD⁺ pool required for PARP1-mediated repair of oxidative damage at telomeres. The authors further report an inverse correlation between donor lymphocyte TL and mitoROS in the corresponding cybrids, and provide preliminary evidence that the K1a-defining ATP6 A177T variant (m.G9055>A) may be enriched in long-telomere individuals.

      We thank the reviewer for this very positive evaluation of our work.

      Weaknesses:

      (1) Statistical support and donor sampling for the central in vivo correlation (Figure 4C).

      The inverse correlation between donor lymphocyte TL and cybrid mitoROS (R<sup>2</sup>=0.794, p=0.007) is the principal in vivo claim of the paper, but it is built on seven donors deliberately selected from the extremes of the Flow-FISH distribution. Sampling at the tails of the outcome variable can substantially inflate apparent correlation strength and significance. I would encourage the authors to (i) explicitly state this sampling structure where the correlation is introduced, (ii) report a leave-one-out sensitivity analysis to confirm the relationship is not driven by one or two donors (Cyb3 and Cyb6 appear to anchor the line), and (iii) where feasible, extend the analysis to additional donors with intermediate TL to test whether the relationship holds across the full distribution. Even a modest expansion (e.g., 4 to 5 additional donors at P25 to P75) would substantially strengthen this central claim.

      We thank the reviewer for this insightful comment. As suggested, we will explicitly describe the sampling structure upon introducing the correlation analysis. Furthermore, we have conducted the requested leave-one-out analysis:

      - removing Cyb3: R<sup>2</sup>=0.730; p=0.0302

      - removing Cyb6: R<sup>2</sup>=0.782; p=0.0194

      - removing Cyb1: R<sup>2</sup>=0.740; p=0.0278

      While we agree that additional donors would enhance the study, further experiments are however currently impossible without new ethical clearances and additional clinical partnerships.

      (2) Reconciling the cybrid CI / TL relationship (Fig 3B) with the absence of a CI / TL relationship in donor lymphocytes (Figure 4A).

      Figure 3B shows a strong correlation between CI activity and TL in cybrids (R<sup>2</sup>=0.87), while Figure 4A shows no correlation between donor CI activity (measured in the same cybrids) and donor lymphocyte TL. The authors acknowledge this, but the manuscript subsequently builds toward a CI-centric model of in vivo TL regulation, which seems to outrun the data. The most internally consistent interpretation is that the cybrid CI phenotype reports a sensitized in vitro response to the acute oxidative stress of the metabolic shift, rather than a steady-state determinant of leukocyte TL. I would suggest reframing the abstract, significance statement, and Discussion to make this distinction clearer. The in vitro CI / NAD⁺ / PARP1 axis is a strong finding on its own, while the in vivo role of CI activity (as opposed to ROS more broadly) is not yet established here. Donor #1's profile (very long lymphocyte TL, low CI activity, severe shortening in cybrids, no telomere inheritance in offspring) is informative in this regard and could be discussed more directly as a case that helps delineate where the cybrid model does and does not recapitulate in vivo biology.

      We acknowledge that our study does not establish the in vivo role of CI activity in TL regulation. Our abstract specifically highlights an in vitro phenomenon: “Under the specific conditions of cybrid formation, which involve a metabolic shift from glycolysis to oxidative phosphorylation, mtDNA variants associated with reduced CI activity induced rapid telomere shortening, …”.

      We are nevertheless happy to revise the text to clearly separate our in vitro results from in vivo biology as requested.

      (3) The K1a / ATP6 A177T inheritance claim.

      The proposal that K1a (and specifically ATP6 A177T) contributes to maternal inheritance of long telomeres is intriguing but currently rests on three pedigrees (one of which, donor #1, does not support the hypothesis) and a chi-square test that does not reach significance (p=0.153, Figure 4F). The supporting evidence is also limited by the fact that platelet-mediated mitochondrial transfer delivers donor mitochondrial proteins, lipids, and residual mtRNA in addition to mtDNA, making it difficult to attribute the cybrid phenotype of donor #6 specifically to the ATP6 A177T variant. I would recommend either: (a) extending the genotyping screen to additional unrelated donors and, if feasible, confirming the effect of ATP6 A177T through an isogenic approach (e.g., mtDNA base editing in a clean background), or (b) softening the relevant statements to "suggestive trend warranting larger studies," and presenting the K1a observation as hypothesis-generating rather than supportive. The Ashkenazi-centenarian connection raised in the Discussion is an excellent direction for follow-up and could be framed accordingly.

      We agree that the evidence for the AT6 A177T inheritance claim remains inconclusive. To clarify, we do not argue that this mitochondrial variant is solely responsible for longer telomeres; indeed, the mtDNA genome of donor #1 suggests otherwise. Furthermore, the phenotypic impact of such mtDNA variants likely depends on nuclear variants in other telomere-related genes (e.g., hTERT or hTR), meaning AT6 A177T may not consistently result in elongated telomeres. Unfortunately, our ethical protocol precludes screening additional unrelated donors. We will revise the text to soften our statements accordingly.

      New references:

      DeBoy EA et al. Familial clonal hematopoiesis in a long telomere syndrome. 2023. N Engl J Med 388, 2422-2433.

      Davidson-Swinton HR et al. Lymphoid malignancy and clonality in the POT1-mediated long telomere syndrome. 2026. Blood 147, 2226-2237.

      Maranzana E et al. Mitochondrial respiratory supercomplex association limits production of reactive oxygen species from complex I. 2013. Antioxid Redox Signal 19, 1469-1480.

    1. eLife Assessment

      This fundamental manuscript describes a key role for the integrated stress response-regulated transcription factor CHOP in regulating liver biology in response to endoplasmic reticulum stress through both the downregulation of transcription factors involved in regulating hepatic identity and altering the capacity for integrated stress response and unfolded protein response signaling to induce protective signaling. The data supporting this model is convincing, but including some additional discussion on the mechanism and importance of the work in the context of the published literature would be helpful to better define the complex importance of CHOP signaling. This work will be of interest to a wide range of biologists interested in liver biology, stress-responsive signaling, and ER stress.

    2. Reviewer #1 (Public review):

      Summary:

      The predominant view on CHOP's functions during ER stress is that it promotes cell death. This is in contrast to a handful of reports in the literature that claim that CHOP is a positive regulator of protein synthesis during chronic ER stress, and therefore is part of the adaptation program to ER stress. These previous studies were performed in tissue culture cells. Velarde and co-authors have used a mouse model of induction of mild ER stress to study the function of CHOP in hepatocytes.

      Major strengths and weaknesses of the methods and results:

      The authors use state-of-the-art mice to manipulate (i) CHOP and (ii) ATF6, a protective factor of ER proteostasis, and address the hepatocyte responses to mild ER stress in vivo and in cultures. Validated gene expression programs are well correlated to liver pathology in the mouse models. This is a very well-done study.

      The authors clearly show that CHOP transitions hepatocytes under mild ER stress to a chronic ISR state, which is phenocopied by ATF6-depleted hepatocytes. So the conclusion that CHOP exacerbates ER stress in hepatocytes during mild ER stress is correct. It is also clear that CHOP targets negatively the transcription of hepatocyte identity genes, which opens a new direction of studies on the function of CHOP in secretory cells in general.

      Conclusion:

      This is a significant study that will benefit different research fields, and specifically studies on proteostasis, as was recently highlighted in Nat. Str. Mol. Biol. by experts in the field.

      To this reviewer, the importance of the study is that it links the function of a transcription factor (CHOP) to stress intensity (mild versus severe) in a physiological experimental model (hepatocyte function and pathology).

    3. Reviewer #2 (Public review):

      The Unfolded protein response (UPR) and related integrated stress response (ISR) are critical signaling systems for cell survival in response to acute stresses. While the UPR directs critical adaptive gene expression, certain chronic stresses switch this pathway towards cell death and disease. An important question concerns the mechanisms by which the UPR switches from being adaptive to maladaptive. Prevailing models focus on the transcription factor CHOP (DDIT3 or GADD153), whose levels are enhanced via the UPR, and extended/amplified amounts of CHOP are suggested to boost death-related gene expression. However, the literature and this manuscript point out a number of observations that do not neatly fit with this model, suggesting that there are still unresolved processes by which CHOP adjusts cell outcomes via the UPR.

      This manuscript features a nice hepatocyte-targeted knockout of CHOP to discern the contribution of CHOP in the transition between adaptive and maladaptive outcomes. The key ideas presented in this study are that CHOP-directed gene expression is focused on protein synthesis, metabolism, and hepatocyte identity. In the progression of the UPR, CHOP expression can lead to resumption of protein synthesis, which can assist in the translation of the UPR-directed transcriptome, which includes ATF6/XBP1-directed genes that aid the processing capacity of the endoplasmic reticulum (ER). However, enhanced nascent protein can further stress the ER. CHOP directs gene expression in both the first phase- acute and second phase-chronic in the UPR, and the pivotal decision lies in the transition between the phases.

      Overall, the manuscript includes some new ideas as well as refinements of earlier ones for CHOP-determination of UPR-directed cell fate. The CHOP-hepatocyte knockout mouse model helps to delineate the different tissue functions of CHOP, which has been a problem for some earlier studies. The manuscript progression of experiments is solid, and experimental design and documentation are rigorous. The manuscript text is largely clear, but there are portions that would benefit from fuller explanations of ideas.

      There are three points of concern. First, the manuscript model (Figure 7) lays out a timeline for the progression of the UPR between two phases. The study is not always clear about the times assayed, and there appears to be a single time point for measurements. Second, there is emphasis on protein synthesis changes in the model. It is true that the literature argues that resumption of protein synthesis concurrent with stress damage (i.e., GADD34-directed gene expression) is a key reason for the potentially debilitating effects of CHOP (e.g., Marciniak et al 2004, Han et al 2013). However, the manuscript does not feature protein synthesis measurements. Inclusion of bulk protein synthesis measurements in the context of this model system would strengthen the study and support for the model. Finally, for this reviewer, some of the most interesting ideas center on CHOP-directed transcription of genes that regulate hepatocyte identity. There is solid evidence for direct CHOP regulation of these genes, but the manuscript does not really develop and test the ramifications of these networks on cell fate during ER stress.

      Reviewer Concerns:

      (1) The abstract packs in a lot of information. The ideas would not be clear to a general reader. Furthermore, the UPR and ISR are referred to in the second-to-last sentence, but not defined earlier in the abstract.

      (2) There are some typos/grammar concerns.

      (3) ATF4 diminished with CHOP-depletion (Figure S2A). What is the mechanism here? Does this complicate the analysis of CHOP-directed gene expression? How does this fit with Figure 6J? The timelines for TM treatment are critical. The authors should more fully explain the time courses in the experiments.

      (4) Figures 2 and 3: There is a discussion on enhanced protein synthesis with loss of CHOP (reduced GADD34 expression). What is the time point - 8 hours TM? Emphasize, explain, and justify time points of experiments here and in later panels. It would strengthen the model with direct measurements of protein synthesis. The authors could include GADD34 protein measurements in these panels. Figure 3 - panel D - some abbreviations are not standard.

      (5) Figure 4: One of the most interesting in the manuscript is the transcription factors downstream of CHOP that are linked with hepatocyte differentiation and metabolism. The manuscript would be bolstered by developing some of these target genes into the Figure 7 transition model.

      (6) Figure 6: The comparison of CHOP and ATF6 target genes is a highlight of the manuscript. The literature on this topic is complex, and there are some suggestions that CHOP can be downstream of ATF6. Furthermore, there were some earlier models by Walter and others about extended induction of Perk (death) vs induction of other UPR sensors (survival) (e.g. PMID: 17991856). It would be helpful in the Discussion to delineate between these models and their critical differences.

    4. Reviewer #3 (Public review):

      In this manuscript, the authors aim to understand the function of the transcription factor CHOP, which is known to promote cell death during severe stress in the ER. The authors note that CHOP is induced during less severe stress, but its functional output is not well understood in these cases. Here, they study the effects of conditional knockouts of CHOP in hepatocytes of mice challenged with chemical inducers of ER stress.

      Tunicamycin (an ER stress inducer) injection leads to the upregulation of CHOP and lipid accumulation in the liver, but no significant cell death in the experiments outlined here. Conditional knockout of CHOP results in a number of differences in the way hepatocytes respond to stress, notably resulting in lower steatosis.

      There are two main findings supported by the data presented here. First, the authors show that CHOP suppresses the expression of ONECUT, a master regulator of hepatocyte differentiation and metabolism, during ER stress. They show by ChIP-seq that CHOP binds to the promoter region of this gene, and by RNA-seq that ONECUT expression is suppressed by ER stress in a CHOP-dependent manner. Many predicted targets of ONECUT1 were also suppressed by ER stress in a CHOP-dependent manner, though they were not bound directly by CHOP. The data support a model where CHOP down-regulates hepatocyte metabolism and identity via regulation of ONECUT1. This is a new and interesting finding, perhaps explaining the steatosis phenotype of livers that accompanies ER stress, although this was not tested directly.

      The second main finding of this paper is that CHOP deletion leads to an interesting assortment of effects on genes related to the ER stress response and integrated stress response (ISR). As expected, based on prior work, CHOP deletion led to more phosphorylation of eIF2alpha (CHOP is known to upregulate the phosphatase for this translation factor). However, unexpectedly, this did not cause increased expression of ATF4 (a transcription factor whose upregulation during stress is dependent on eIF2alpha phosphorylation) and its downstream targets; in fact, CHOP deletion had the opposite effect on these. In other words, CHOP seems to both turn off the initiating signal for the ISR (namely, eIF2alpha phosphorylation) and also promote the downstream signaling events that rely on this initiating signal. It makes sense that cells would do this, as restoring translation would be important for realizing the effects of the massive changes in gene expression initiated by ER stress, and yet this would exacerbate stress in the short term, so it would be counterproductive to also turn off the entire stress-regulated program. Having a factor (perhaps CHOP) that coordinates these two events makes sense. It will be interesting in future work to understand the mechanisms behind this regulation.

      Finally, CHOP deletion led to less activity of other aspects of the ER stress response, notably IRE1 (determined through measurement of XBP1 splicing and RIDD of Bloc1s1). This is explained by the continued phosphorylation of eIF2alpha in these knockouts, as the continued attenuation of translation would lessen the burden of misfolded proteins in the ER. Somewhat confusingly, the same pattern is not seen in downstream targets of XBP1. Less splicing, coupled with perhaps less translation of the spliced mRNA, should result in less active transcription factor and lower expression of its target genes in the CHOP KO. This is not observed in Figure 2, although the more global gene expression analysis suggests that all stress-dependent gene expression changes were weaker in the CHOP KO livers.

      The authors characterize the effects of CHOP, promoting restoration of protein synthesis and the accompanying exacerbation of stress while preserving the signaling that should relieve ER stress, as a switch from an acute to chronic phase of ER stress. This is mirrored in their analysis of ATF6 in a similar series of experiments. Although this is an interesting framework for thinking about the stress response, whether CHOP is the key factor or a supporting actor in regulating this transition will require a better understanding of the mechanisms involved.

    5. Author response:

      We thank the reviewers for their assessment of our work and their comments. We are grateful for their evaluation of our findings as fundamental and convincingly supported, and for their appreciation of the relative scope of this manuscript and of future work. The most direct requests for new experimental data are from reviewer #2, who asks for direct assessment of the effects of CHOP deletion on expression of GADD34 and on protein synthesis. We agree that these are important experiments to conduct for the revision.

      The reviewers requested more clarity on the experimental logic of the paper and on the place of our findings in the broader context of ER stress signaling, which we will be happy to provide in a revised manuscript. These revisions will include a more explicit consideration of how the regulation of metabolic genes by CHOP contributes to its effects in the liver independently of its role in regulating eIF2a dephosphorylation.

      In particular, there were concerns about the logic of the time points chosen that we feel are important to also address here. For analysis of ChopHKO animals, all experiments were carried out 8 hours after ER stress challenge. This is because, as we show in Fig. 1B and also in our previous paper on CHOP (1), this is the time point at which CHOP expression is at its maximum. Thereafter, hepatocytes become heterogeneous with respect to whether they do or do not express CHOP. This is an interesting finding because it suggests that CHOP is part of a cellular switch, and potentially even an effector of that switch—a point currently raised in the Discussion but worth further highlighting in a revision. At the practical level, it means that discerning the contribution of CHOP to ER stress signaling and adaptation at subsequent time points will require sophisticated single cell analyses that can discriminate cells that express CHOP from cells that do not, which are an important future direction.

      In contrast, for Atf6aHKO animals, all experiments were carried out 48 hours after ER stress challenge. As we have previously shown (2), at short time points after a stress challenge, such as 8 hours, there is very little difference in ER stress signaling between wild-type animals and those lacking ATF6a. The reason for this lack of distinction is that the major targets of ATF6a are ER chaperones and the like. Because adaptation to ER stress in the early phases of the response depends more on non-transcriptional mechanisms such as inhibition of protein synthesis and IRE1-dependent mRNA decay (RIDD), the failure to fully upregulate ATF6a targets is initially of little consequence. It is only at later time points when wild-type animals restore ER homeostasis and largely silence ER stress signaling. In contrast, at these same later points, animals lacking ATF6a show evidence of persistent ER stress, most notably in the form of persistent Xbp1 mRNA splicing and profound suppression of metabolic genes. Although the 8 hour time point for experiments in ChopHKO animals differs from the 48 hour time point for Atf6aHKO animals, the two lines of experimentation are united by the persistence of ongoing ER stress and of ISR signaling despite diminished eIF2a phosphorylation at the points when the presence of CHOP or the absence of ATF6a are of the greatest impact. A revised manuscript will present this logic more clearly.

      References

      (1)  Liu K, et al., EMBO Reports 25, 228 (2024)

      (2)  Rutkowski DT, et al., Dev. Cell 15, 829 (2008)

    1. eLife Assessment

      This important study combines behavioural analysis, voltage imaging and electrophysiology to advance our understanding of muscle coordination at the cell-to-cell level, in Caenorhabditis elegans. The evidence supporting the conclusions is convincing; however, the use of correlation in some aspects of data interpretation is a relative weakness. The technically sophisticated optogenetic voltage clamp approach introduced here can be applied to other small, transparent animals, making these findings of broad interest to researchers studying electrical coupling between cells or utilising optical electrophysiology techniques.

    2. Reviewer #1 (Public review):

      Summary:

      This study aims to reveal the contribution of individual gap junction proteins to the signal transmission and connectivity of living C. elegans animals in a completely non-invasive way through all-optical electrophysiology. The authors achieve this by simultaneous expression of bipoles, an excitatory/inhibitory light-activated actuator and Quasar2, a genetically encoded voltage dye. With this study, the authors extend their previous efforts to leverage the strength of optogenetic neurophysiology and set a new standard in this domain. In addition, they adapted their established methods to perform cell-specific optogenetic voltage clamp and revealed changes in gap junction connectivity. They also find that increasing excitability in innexin mutants is indicative of a reduction in gap-junction connectivity and current leaks.

      Strengths:

      This is an extremely strong manuscript, a technical feat and tour de force to infer junctional coupling through all-optical electrophysiology. The establishment of the voltage clamp method is powerful and allows researchers to obtain not only tight control over voltage signals but also permits the investigation of gap junction function in response to positive and negative voltage steps in a completely non-invasive fashion. This will be a new paradigm for investigating muscle electrophysiology in future.

      Weaknesses:

      This is a strong pioneering study, and I found very few technical weaknesses. The correlation quantification is relatively weak to establish connective causality, as a shared upstream input may lead to a similar perceived correlation. This is especially concerning for an average lag time of ~0, and the authors may want to investigate if there is unchanged connectivity in an unc-31 or unc-13 mutant. Conceptually, the local connectivity is scaled to account for behaviour: future studies may wish to perform this method on moving animals, and in specific neuronal populations, where a non-invasive optogenetic voltage clamp method will truly shine.

    3. Reviewer #2 (Public review):

      Summary:

      This technically sophisticated study combines behavioral analysis, voltage imaging, electrophysiology, and a newly developed cell-specific optogenetic voltage clamp (cOVC) approach to investigate gap-junction (GJ)-mediated coupling in C. elegans body-wall muscle cells. The work explores the coordination of muscle cells and systems physiology and introduces a method with potential utility beyond the nematode system studied.

      Strengths:

      The main strength of the work is the development and application of the cOVC method. This approach enables minimally invasive in vivo assessment of cell-to-cell electrical coupling in intact animals. This technique represents a meaningful advance over traditional electrophysiological techniques that require dissection or cell isolation.

      With respect to the GJ biology and function, the authors support their conclusions by integrating additional independent experimental approaches. Findings from behavioural/locomotion assays, voltage imaging, patch-clamp recordings, and cOVC measurements are generally consistent, particularly for unc-9 mutants, which show reduced synchronization of muscle cells, impaired electrical coupling, and severe locomotor defects.

      The gain-of-function experiment using murine Cx36 suggests that more or less electrical coupling can disrupt (normal) locomotion.

      Weaknesses:

      The main issue of this otherwise excellent manuscript relates to interpretation rather than experimental quality. Throughout the manuscript, increased correlation is often interpreted as evidence of increased electrical coupling. Are correlation, synchrony, and conductance equivalent measures? If not, how would this affect these correlations? Furthermore, could broader action potentials and altered excitability also increase correlation values? This concern could be addressed through a discussion of this limitation.

      Similarly, the proposed mechanism that reduced GJ coupling increases excitability through reduced leak currents is plausible but not directly demonstrated. Are alternative explanations, e.g., compensatory changes in ion-channel expression or gap-junction composition, possible? These could also be considered to improve the balance of this work.

      The conclusions regarding Cx36 overexpression would also benefit from more cautious wording, as developmental or localization effects have not been excluded.

      However, the experimental dataset is very strong. In my opinion, no major additional studies are needed. Direct analysis of compensatory changes in innexin expression or localization could strengthen the interpretation of the proposed mechanism. Overall, the study is of high technical quality, contains a notable methodological advance, and provides important insights into muscle synchronization and GJ biology.

    1. eLife Assessment

      This important study describes how selective lesions of key cortical and subcortical motor areas affect reaching actions in macaque monkeys. The results will be of interest to both basic and clinical researchers studying the neural control of movement. Kinematic analysis of movement quality is solid but could be improved by considering other metrics, especially those that relate to grasping. Evidence for the general claims related to the role of specific motor areas is incomplete because the lesions did not fully eliminate any single area while simultaneously involving neighbouring areas.

    2. Reviewer #1 (Public review):

      Summary:

      This is a very interesting and well-done study of the effects of selective lesions to the sensorimotor cortex and the red nucleus on control of upper limb movements. The findings that the red nucleus may subserve recovery of upper limb motor function after cortical lesions in macaques and the different motor functions of different cortical sensorimotor areas are significant findings of considerable interest to sensorimotor neuroscientists, neurologists and neurosurgeons. The methods are mostly excellent, but there are some questions about the use of endothelin lesions in cortical areas and the use of trajectory variability as a marker of movement quality and fine motor control. Furthermore, it is questionable that increased trajectory variability in reaching a target reflects reduced movement quality, reduced ability to independently control muscles, and is a proximal analog of reduced dexterity.

      Strengths:

      The rationale that rubrospinal projections onto spinal neurons may subserve the good recovery of upper limb movements observed after lesions of sensorimotor cortex is compelling. The methods involving complete lesions of the red nucleus followed by recovery prior to lesions affecting various sensorimotor cortical areas are a strength. The excellent interpretations offered in the Discussion section are also a strength.

      Weaknesses:

      There are weaknesses in the Methods, including:

      (1) no information on dimensions of the cup containing the food reward or types of food rewards,

      (2) recording 3D hand movements with a single camera,

      (3) cortical endothelin lesions were not very precise,

      (4) the use of trajectory variability as a measure of movement quality and reduced ability to independently control muscles.

      Some interpretations presented in the Discussion are not well supported. The discussion related to movement quality should be modified to focus on trajectory variability. The suggestion that rubrospinal projections onto motor neurons are apparently irreplaceable is not well justified because one monkey receiving a complete red nucleus lesion showed nearly full recovery of maximum movement speed, while the other monkey did not. The nearly full recovery of one monkey was probably due to new corticospinal connections onto motor neurons, whereas it is possible that the other monkey would have recovered better given more time before the 2nd lesion to cortical areas.

    3. Reviewer #2 (Public review):

      Summary:

      This study made selective lesions in motor cortical subregions and the magnocellular red nucleus in nine macaque monkeys, and evaluated reaching and grasping movements using maximum speed and trajectory variability. The results suggest that damage to the posterior old primary motor cortex (M1) was mainly associated with reduced maximum speed, whereas damage to the new M1 was mainly associated with increased trajectory variability. Damage to the anterior old M1 did not clearly add further impairment. Lesions of the magnocellular red nucleus (RNm) alone mainly reduced reaching speed, but recovery after subsequent cortical lesions was worse than after cortical lesions alone, suggesting that the rubrospinal pathway may be important for compensation after cortical damage in monkeys. Overall, this is a valuable study that examines differences among M1 subregions using selective lesions in macaques.

      Strengths:

      (1) This study tackles an important question. It attempts to decompose the diverse upper-limb impairments after stroke into the effects of different primary motor cortex subregions.

      (2) Another strength is that the lesions and behavioral impairments were evaluated quantitatively. The use of nine macaque monkeys with different lesion patterns, together with quantitative behavioral evaluation, provides a rare and valuable dataset. The authors also followed recovery using quantitative behavioral measures such as maximum speed and trajectory variability.

      (3) The inclusion of RNm lesions is also valuable, as it revisits the classic question raised by Lawrence and Kuypers (1968) of how brainstem descending pathways contribute to recovery after cortical motor damage.

      Weaknesses:

      (1) The main limitation is that the contribution of each cortical lesion is sometimes interpreted from largely qualitative comparisons. Because the lesion extent was not always limited to the intended region, it is difficult to fully separate the independent contribution of each subregion. Some conclusions are also based on comparisons between a small number of animals. The dataset itself is valuable, and the manuscript would be strengthened by presenting these conclusions more cautiously and explicitly acknowledging this limitation.

      (2) Because the behavioral evaluation is quantitative, it would be helpful to show the relationship between lesion size and behavioral impairment more quantitatively. For example, rank correlations between the lesion size of each cortical region and behavioral measures could help readers evaluate whether the type and size of lesion are related to behavioral impairment.

      (3) The discussion of area 4s could be further developed. The authors suggest that this region may have a different role, but the specific hypothesis is not fully clear. There has also been skepticism in the previous literature about area 4s, for example, Meyers et al. (1954), and this broader background could be discussed. (Meyers R, Knott JR, Skultety FM, Imler R (1954) On the Question as to the Existence of a "4s" Suppressor Mechanism. Journal of Neurosurgery 11:7-23.)

    4. Reviewer #3 (Public review):

      Summary:

      In this article, the authors performed targeted lesions in cortical areas involved in forelimb control of rhesus macaques. Using a reaching task with kinematic tracking, they compared kinematic variability (as a proxy for dexterity) and reaching speed (as a proxy for strength) before and after cortical (n=7) and magnocellular red nucleus (RNm) (n=2) lesion. Changes in these movement metrics were related to the location and extent of the lesions, reconstructed from histology. The authors report that lesions with a large component in New M1 had a pronounced effect on kinematic variability, whereas lesions with a large component in posterior Old M1 primarily affected reaching speed. Lesions of the RNm were performed in two animals approximately seven to eight weeks before the cortical lesion. By themselves, RNm lesions produced a significant but small reduction in reach speed. They also magnified the effect of the subsequent cortical lesions.

      Strengths:

      (1) For non-human primate (NHP) research, this is a large cohort.

      (2) The behavioural analyses are clear and precise.

      (3) The additional red nucleus lesions in two monkeys provide unique complementary information.

      Weaknesses:

      (1) Description of injuries. As described and reported in the current result section and figures, readers do not have a clear understanding of the lesion extent and location.

      (2) Lack of formal correlative analyses between lesion characteristics and behaviour. Currently, it seems that the conclusions are based on impressions between some aspects of the lesions and precise kinematic measures.

      (3) The data will be of interest to a large community of researchers working on brain injury and stroke recovery, as well as cortical motor control. There are, however, some major methodological issues in the current version of the manuscript that prevent a clear evaluation of the findings and potential contribution of the work in the field. These issues need to be addressed to support the conclusions.

    1. eLife Assessment

      The framework of this potentially important study – with the integration of multiple levels of analysis, glymphatic MRI, transcriptomics, functional MRI, and public amyloid maps, in one framework – is clever. The assertion that regional amyloid vulnerability may depend not just on neural activity alone, but on whether clearance is appropriately matched to activity, is an interesting and novel concept. However, the chosen approach to imaging glymphatic clearance relies on indirect inferences from a small subgroup. In its current form, the main conclusions of this study are therefore incompletely supported.

    2. Reviewer #1 (Public review):

      Summary:

      Regional differences in the brain's waste-clearance system may interact with neural activity to influence where amyloid-B accumulates. Using intrathecal GBCA administration to produce "Glymphatic MRI" in 96 subjects, the authors mapped cortical glymphatic influx and clearance and found distinct spatial patterns, with transcriptomic analyses linking better glymphatic function to neuronal cell types (through genes). In a subgroup with resting-state fMRI, regions with stronger resting-state activation generally showed higher contrast clearance, indicating a positive coupling between these processes. Notably, cortical regions where neural activity and glymphatic clearance were mismatched showed greater amyloid-β burden in a separate, publicly available PiB-PET dataset, suggesting that activity-clearance decoupling may contribute to regional vulnerability and neurodegeneration.

      Strengths:

      This is a rare and valuable dataset. Intrathecal contrast injection in ~100 subjects is quite a remarkable accomplishment alone, but the addition of resting-state fMRI, a correlative PiB cohort, and gene-expression pattern data is impressive.

      Weaknesses:

      This is a cross-sectional study, and we can't determine whether neural activity drives glymphatic clearance, whether glymphatic dysfunction alters neural activity, or whether both are shaped by a third factor. Language describing "flow", "influx", and "clearance" could be made more specific so the reader can more easily follow the methodological approach.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, Li et al. investigated the relationships among regional cortical tracer dynamics following intrathecal gadolinium administration, neural activity, and amyloid-β deposition in humans. Using serial MRI acquisitions after intrathecal gadodiamide administration in 96 participants, the authors characterized regional signal enhancement and clearance patterns across the human cortex. They integrated these imaging measures with transcriptomic data (Allen Human Brain Atlas), resting-state fMRI outcomes, and an external amyloid PET dataset. The authors report that regions with more efficient tracer clearance are enriched for genes related to synaptic organization and neuronal cell types, that tracer clearance patterns are in parts spatially coupled to spontaneous neural activity, and that regional mismatch between neural activity and tracer clearance is associated with increased amyloid burden according to the PET dataset.

      Strengths:

      The study addresses an important and very timely question about the interaction among neural activity, cerebrospinal fluid dynamics (waste clearance), and regional vulnerability to neurodegeneration. Integrating serial post-contrast MRI, transcriptomics, resting-state fMRI, and amyloid imaging is ambitious and conceptually very interesting. The spatial characterization of cortical tracer dynamics is potentially valuable for the field, particularly given the increasing interest in human glymphatic imaging approaches and intrathecal contrast MRI, which provides an opportunity to assess CSF tracer dynamics without confounding tracer signal from the blood. The imaging preprocessing pipeline includes normalization of regional cortical signal intensity to a reference region within each session before calculation of longitudinal percentage change, which helps reduce inter-session variability within individuals for conventional T1-weighted imaging. The transcriptomic analyses linking tracer dynamics to neuronal and synaptic gene expression patterns are also interesting. In addition, the manuscript addresses recent literature on neurovascular coupling, glymphatic function, and amyloid vulnerability.

      Weaknesses:

      Several issues limit the strength of the conclusions. One concern relates to the interpretation of repeated post-intrathecal contrast MRI measurements as direct indicators of glymphatic influx and clearance. The approach presented by the authors measures regional signal changes following intrathecal gadodiamide administration, but does not directly visualize paravascular flow or establish that the observed signal dynamics specifically reflect glymphatic transport mechanisms. Although it is widely accepted that CSF influx occurs primarily along periarterial spaces as part of the glymphatic system, and the terminology "glymphatic MRI" is increasingly used in the literature, the physiological processes contributing to delayed parenchymal enhancement, including CSF-interstitial exchange mediated by convective bulk flow and/or extracellular diffusion, as well as transient and, in the case of linear gadolinium agents, even long-term tracer retention remain incompletely resolved. Importantly, tracer kinetics may not directly reflect interstitial fluid kinetics, as solute transport may also be influenced by compartmental and extracellular barriers, diffusion constraints, and tissue retention effects. As currently written, several sections of the manuscript appear to overstate what can be directly inferred from the imaging data. This issue may be particularly relevant given the intrathecal use of gadodiamide (Omniscan), a linear gadolinium-based contrast agent with known long-lasting tissue retention due to lower kinetic stability compared to macrocyclic agents. Sustained signal at later imaging time points may therefore not only reflect impaired glymphatic clearance dynamics may also be influenced by tissue retention of contrast material, particularly in the context of neurological disease. In addition, the participant cohort is heterogeneous and includes individuals with neuroinflammatory and neurodegenerative diseases, peripheral neuropathy, and motor neuron disease. Although the authors argue that the spatial tracer patterns are relatively preserved across neurodegenerative groups, this heterogeneity complicates interpretation of imaging data and raises the possibility that disease-related factors and altered tracer-tissue interactions contribute to the observed effects. Thus, the rationale for interpreting a greater tracer signal at 39h as evidence of impaired glymphatic clearance should be explained more carefully, particularly given the highly heterogeneous patient population.

      In addition, the analyses linking spontaneous neural activity and tracer clearance are based on a very small rs-fMRI subgroup (n = 15), limiting the generalizability. The interpretation of the "mismatch" analysis also requires caution. The mismatch index was computed from z-scored fALFF and tracer clearance and is subsequently associated with amyloid burden derived from the external PET dataset rather than from the studied participants themselves. Therefore, the observed spatial associations should be interpreted with greater caution rather than as evidence for a direct mechanistic relationship. The cross-sectional nature of the analyses also limits conclusions regarding the directionality and temporal sequence of the relationships between neural activity, tracer dynamics, and amyloid burden. Several statements in the Discussion currently imply stronger causal or biological conclusions than are directly supported by the data.

      Despite these limitations, the study presents an interesting dataset and proposes a framework for understanding regional vulnerability to protein accumulation in neurodegeneration. This work hopefully motivates further investigation into the important relationships among neural activity, CSF dynamics, and neurodegeneration in humans.

    4. Reviewer #3 (Public review):

      This manuscript addresses an interesting and timely question: whether regional glymphatic clearance in the human cortex is spatially coupled to neural activity and whether a mismatch between activity and clearance may help explain regional vulnerability to amyloid-β deposition. The authors use intrathecal gadolinium-based glymphatic MRI in 96 participants, derive cortical influx and clearance maps, integrate these with Allen Human Brain Atlas transcriptomic data, and then relate regional clearance to resting-state fMRI measures in a smaller subgroup. They further compare the resulting activity-clearance mismatch map with an open-source ¹¹C-PiB amyloid PET dataset. The overall concept is attractive because it attempts to connect glymphatic physiology, neuronal activity, and proteopathy at the regional level of the human brain, an important and understudied area.

      The main strength of the study is the use of direct intrathecal contrast-enhanced MRI to generate cortical maps of glymphatic tracer dynamics. This is a technically demanding approach and provides a richer spatial readout than indirect MRI proxies of glymphatic function. The authors show that the cortical tracer signal increases from 4.5 h to 15 h and then decreases by 39 h, allowing them to interpret the early signal as reflecting influx and the persistent signal at 39 h as impaired clearance. They further identify regional patterns, with faster influx in medial prefrontal/insular areas and slower clearance in dorsal prefrontal and parietal surface regions. The analysis is visually clear, and the use of cortical gradients is a useful way to reduce complex regional data into interpretable spatial axes.

      The multimodal integration is also interesting. The transcriptomic analysis suggests that regions with faster glymphatic clearance are enriched for synaptic organisation and neuronal activity-related pathways, while regions with slower clearance show enrichment for metabolic and mitochondrial pathways. The cell-type enrichment analysis further implicates excitatory and inhibitory neurons, oligodendrocyte lineage cells, microglia and, to a lesser extent, astrocytes. This provides a plausible biological bridge between regional neural activity and clearance function, and the sensitivity analysis using ReHo in addition to fALFF is a useful robustness check.

      However, the manuscript should be more careful in its causal interpretation. The study is cross-sectional and largely correlative in space. The finding that regions with higher spontaneous neural activity tend to show better glymphatic clearance is intriguing, but it does not establish that neural activity drives clearance in these participants. Conversely, it remains possible that better tissue integrity, vascular function, CSF access, cortical geometry, vascular density, or disease composition jointly influence both fMRI measures and tracer clearance. The authors do acknowledge some of these limitations, but the abstract and discussion should more consistently frame the findings as associations rather than evidence of an activity-clearance mechanism in humans.

      The most important limitation is the small size of the fMRI subgroup. Although the whole glymphatic MRI cohort includes 96 participants, the key activity-clearance analysis is based on only 15 individuals, including 11 with peripheral neuropathy and 4 with motor neuron disease. This is a very small and clinically heterogeneous sample on which to build a central conclusion about regional neural activity and glymphatic clearance. The authors show that the 39 h PC map in the fMRI subgroup resembles the whole-cohort map, which is helpful, but this does not address whether the fALFF-clearance relationship is robust at the individual level. The paper would be strengthened by reporting subject-level stability, leave-one-out analyses, and whether the association persists after excluding the four motor neuron disease cases.

      A second major concern is the interpretation of the amyloid analysis. The ¹¹C-PiB map is derived from an external open-source Alzheimer's disease dataset, not from the same participants who underwent glymphatic MRI and fMRI. Therefore, the association between activity-clearance mismatch and amyloid burden is a spatial correspondence across group-average maps, not an individual-level relationship. This is valuable for hypothesis generation, but should not be presented as evidence that a mismatch in the present cohort predicts amyloid deposition. The authors should clearly state that this analysis tests whether mismatch regions overlap with known amyloid-prone cortical regions, rather than directly linking mismatch to amyloidosis in individual participants.

      The definition of "mismatch" also needs clarification. The text defines the mismatch index as the negative absolute difference between z-fALFF and z-39h PC, and states that higher scores indicate greater mismatch. Because the index is negative, values closer to zero would normally indicate a smaller absolute difference rather than a greater mismatch. This should be checked carefully and corrected if necessary. More broadly, because a higher 39 h PC indicates worse clearance, the interpretation of match and mismatch categories is not intuitive. The authors should provide a clearer schematic and ensure that the mathematical definition, biological interpretation and figure labelling are fully aligned.

      Several technical confounds require more attention. Intrathecal gadolinium MRI is influenced by CSF dynamics, posture, sleep, circadian timing, renal clearance, age, intracranial pathology, and potentially diagnosis-specific differences. The authors acquired scans at fixed time points and noted that patients slept as usual, but individual sleep duration, sleep quality, posture, and daytime activity were not objectively measured. Given that the central claim concerns glymphatic clearance, these are not minor confounders. The authors should consider adjusting for age, sex, diagnosis, vascular risk factors, and relevant clinical variables where possible, and be more explicit about how heterogeneous disease indications may influence cortical tracer kinetics.

      The statistics are generally good. However, many correlations are performed across 400 cortical parcels, which are not independent biological samples. The paper would benefit from clearer separation between participant-level inference and region-level spatial inference. For example, the fALFF-clearance and mismatch-amyloid analyses are regional map correlations, not correlations across individuals. This should be clearly stated throughout. The authors should also report effect sizes and confidence intervals more consistently, and explain how multiple comparisons were controlled across transcriptomic, cell-type, fMRI, ReHo and amyloid analyses.

      The transcriptomic analysis is useful but should be presented as indirect. AHBA data come from six post-mortem brains; only the left hemisphere was used, and the donors were healthy and younger than the clinical cohort. Therefore, these data capture intrinsic regional gene-expression patterns rather than disease-state expression in the same individuals. The authors should avoid implying that the transcriptomic findings directly explain glymphatic function in their participants. The current discussion partly acknowledges this, but the framing in the abstract and results could be more cautious.

      There are also several points of presentation that should be improved. The manuscript should consistently distinguish glymphatic influx, glymphatic clearance, CSF tracer retention, and waste clearance. A 39 h residual gadolinium signal is a useful proxy for delayed clearance, but it is not the same as direct measurement of amyloid or tau clearance. The language around "waste clearance" and "amyloidosis" should therefore be precise. The authors should also clarity whether "higher clearance" corresponds to lower 39 h PC across all analyses, as this inversion is easy for readers to misinterpret.

    5. Author response:

      We would like to express our deepest gratitude to the Editors and Reviewers for their highly rigorous and constructive evaluation of our manuscript. We are greatly encouraged by the recognition of our study’s ambition, the unique value of the in vivo intrathecal contrast MRI dataset, and the conceptual novelty of linking macroscopic glymphatic physiology with neural activity and regional proteopathy.

      We fully agree with the thoughtful limitations and methodological concerns raised in the eLife Assessment and the Public Reviews. In our upcoming revised manuscript, we are implementing a comprehensive set of revisions to address these points. Specifically, our planned revisions focus on the following key areas:

      - Tempering Causal Interpretations: We agree that our cross-sectional design precludes definitive causal inferences. We are systematically revising the manuscript to soften causal language (e.g., replacing "drives" with "is spatially associated with"). We will explicitly frame our findings as macroscopic spatial associations and discuss the potential influence of joint physiological confounders.

      - Tightening Terminology and Imaging Physics: We are refining our terminology to more accurately reflect our MRI measurements. We will replace assertive terms like "direct glymphatic flow" with precise descriptors such as "imaging proxies for tracer enhancement and retention." Furthermore, we are expanding the Limitations section to explicitly acknowledge the confounding effects of Partial Volume Averaging (PVE), systemic tracer redistribution, and renal clearance kinetics.

      - Conducting Supplementary Imaging & Robustness Analyses: To address concerns regarding cohort heterogeneity and the sample size of the rs-fMRI subgroup (n=15), we are performing a series of rigorous supplementary analyses. This includes conducting sensitivity analyses (e.g., excluding the motor neuron disease subgroup) and applying leave-one-out cross-validation to rigorously assess the subject-level stability and robustness of the spatial coupling between neural activity and tracer clearance.

      - Clarifying the Conceptual Model and "Mismatch" Index: To improve readability, we are moving the anatomical definitions of the cortical gradients directly into the Results section. Additionally, we are introducing schematic diagram to intuitively explain the mathematical formulation and biological interpretation of the "activity-clearance mismatch" index.

      - Re-framing External Dataset Analyses: We are carefully re-framing the interpretations of the Allen Human Brain Atlas (AHBA) transcriptomic data and the external PiB-PET amyloid dataset, emphasizing that these reflect spatial correspondences of intrinsic regional vulnerability across groups, rather than individual-level direct interactions.

      We believe these revisions will significantly enhance the scientific rigor, clarity, and precision of our study.

    1. Author response:

      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public Review):

      This work provides a new dataset of 71,688 images of different ape species across a variety of environmental and behavioral conditions, along with pose annotations per image. The authors demonstrate the value of their dataset by training pose estimation networks (HRNet-W48) on both their own dataset and other primate datasets (OpenMonkeyPose for monkeys, COCO for humans), ultimately showing that the model trained on their dataset had the best performance (performance measured by PCK and AUC). In addition to their ablation studies where they train pose estimation models with either specific species removed or a certain percentage of the images removed, they provide solid evidence that their large, specialized dataset is uniquely positioned to aid in the task of pose estimation for ape species.

      The diversity and size of the dataset make it particularly useful, as it covers a wide range of ape species and poses, making it particularly suitable for training off-the-shelf pose estimation networks or for contributing to the training of a large foundational pose estimation model. In conjunction with new tools focused on extracting behavioral dynamics from pose, this dataset can be especially useful in understanding the basis of ape behaviors using pose.

      We thank the reviewer for the kind comments.

      Since the dataset provided is the first large, public dataset of its kind exclusively for ape species, more details should be provided on how the data were annotated, as well as summaries of the dataset statistics. In addition, the authors should provide the full list of hyperparameters for each model that was used for evaluation (e.g., mmpose config files, textual descriptions of augmentation/optimization parameters).

      We have added more details on the annotation process and have included the list of instructions sent to the annotators. We have also included mmpose configs with the code provided. The following files include the relevant details:

      File including the list of instructions sent to the annotators:

      OpenMonkeyWild Photograph Rubric.pdf

      Mmpose configs:

      i) TopDownOAPDataset.py

      ii) animal_oap_dataset.py

      iii) init.py

      iv) hrnet_w48_oap_256x192_full.py

      Anaconda environment files:

      i) OpenApePose.yml

      ii) requirements.txt

      Overall this work is a terrific contribution to the field and is likely to have a significant impact on both computer vision and animal behavior.

      Strengths:

      Open source dataset with excellent annotations on the format, as well as example code provided for working with it.

      Properties of the dataset are mostly well described.

      Comparison to pose estimation models trained on humans vs monkeys, finding that models trained on human data generalized better to apes than the ones trained on monkeys, in accordance with phylogenetic similarity. This provides evidence for an important consideration in the field: how well can we expect pose estimation models to generalize to new species when using data from closely or distantly related ones?

      Sample efficiency experiments reflect an important property of pose estimation systems, which indicates how much data would be necessary to generate similar datasets in other species, as well as how much data may be required for fine-tuning these types of models (also characterized via ablation experiments where some species are left out).

      The sample efficiency experiments also reveal important insights about scaling properties of different model architectures, finding that HRNet saturates in performance improvements as a function of dataset size sooner than other architectures like CPMs (even though HRNets still perform better overall).

      We thank the reviewer for the kind comments.

      Weaknesses:

      More details on training hyperparameters used (preferably full config if trained via mmpose).

      We have now included mmpose configs and anaconda environment files that allow researchers to use the dataset with specific versions of mmpose and other packages we trained our models with. The list of files is provided above.

      Should include dataset datasheet, as described in Gebru et al 2021 (arXiv:1803.09010).

      We have included a datasheet for our dataset in the appendix lines 621-764.

      Should include crowdsourced annotation datasheet, as described in Diaz et al 2022 (arXiv:2206.08931). Alternatively, the specific instructions that were provided to Hive/annotators would be highly relevant to convey what annotation protocols were employed here.

      We have included the list of instructions sent to the Hive annotators in the supplementary materials. File: OpenMonkeyWild Photograph Rubric.pdf

      Should include model cards, as described in Mitchell et al (arXiv:1810.03993).

      We have included a model card for the included model in the results section line 359. See Author response image 1:

      Author response image 1.

      It would be useful to include more information on the source of the data as they are collected from many different sites and from many different individuals, some of which may introduce structural biases such as lighting conditions due to geography and time of year.

      We agree that the source could introduce structural biases. This is why we included images from so many different sources and captured images at different times from the same source—in hopes that a large variety of background and lighting conditions are represented. However, doing so limits our ability to document each source background and lighting condition separately.

      Is there a reason not to use OKS? This incorporates several factors such as landmark visibility, scale, and landmark type-specific annotation variability as in Ronchi & Perona 2017 (arXiv:1707.05388). The latter (variability) could use the human pose values (for landmarks types that are shared), the least variable keypoint class in humans (eyes) as a conservative estimate of accuracy, or leverage a unique aspect of this work (crowdsourced annotations) which affords the ability to estimate these values empirically.

      The focus of this work is on overall keypoint localization accuracy and hence we wanted a metric that is easy to interpret and implement, in this case we made use of PCK (Percentage of Correct Keypoints). PCK is a simple and widely used metric that measures the percentage of correctly localized keypoints within a certain distance threshold from their corresponding groundtruth keypoints.

      A reporting of the scales present in the dataset would be useful (e.g., histogram of unnormalized bounding boxes) and would align well with existing pose dataset papers such as MS-COCO (arXiv:1405.0312) which reports the distribution of instance sizes and instance density per image.

      We have now included a histogram of unnormalized bounding boxes in the manuscript, see Author response image 2:

      Author response image 2.

      Reviewer #2 (Public Review):

      The authors present the OpenApePose database constituting a collection of over 70000 ape images which will be important for many applications within primatology and the behavioural sciences. The authors have also rigorously tested the utility of this database in comparison to available Pose image databases for monkeys and humans to clearly demonstrate its solid potential.

      We thank the reviewer for the kind comments.

      However, the variation in the database with regards to individuals, background, source/setting is not clearly articulated and would be beneficial information for those wishing to make use of this resource in the future. At present, there is also a lack of clarity as to how this image database can be extrapolated to aid video data analyses which would be highly beneficial as well.

      I have two major concerns with regard to the manuscript as it currently stands which I think if addressed would aid the clarity and utility of this database for readers.

      (1) Human annotators are mentioned as doing the 16 landmarks manually for all images but there is no assessment of inter-observer reliability or the such. I think something to this end is currently missing, along with how many annotators there were. This will be essential for others to know who may want to use this database in the future.

      We thank the reviewer for pointing this out. Inter-observer reliability is important for ensuring the quality of the annotations. We first used Amazon MTurk to crowd source annotations and found that the inter-observer reliability and the annotation quality was poor. This was the reason for choosing a commercial service such as Hive AI. As the crowd sourcing and quality control are managed by Hive through their internal procedures, we do not have access to data that can allow us to assess inter-observer reliability. However, the annotation quality was assessed by first author ND through manual inspections of the annotations visualized on all of the images the database. Additionally, our ablation experiments with high out of sample performances further vaildate the quality of the annotations.

      Relevant to this comment, in your description of the database, a table or such could be included, providing the number of images from each source/setting per species and/or number of individuals. Something to give a brief overview of the variation beyond species. (subspecies would also be of benefit for example).

      Our goal was to obtain as many images as possible from the most commonly studied ape species. In order to ensure a large enough database, we focused only on the species and combined images from as many sources as possible to reach our goal of ~10,000 images per species. With the wide range of people involved in obtaining the images, we could not ensure that all the photographers had the necessary expertise to differentiate individuals and subspecies of the subjects they were photographing. We could only ensure that the right species was being photographed. Hence, we cannot include more detailed information.

      (2) You mention around line 195 that you used a specific function for splitting up the dataset into training, validation, and test but there is no information given as to whether this was simply random or if an attempt to balance across species, individuals, background/source was made. I would actually think that a balanced approach would be more appropriate/useful here so whether or not this was done, and the reasoning behind that must be justified.

      This is especially relevant given that in one test you report balancing across species (for the sample size subsampling procedure).

      We created the training set to reflect the species composition of the whole dataset, but used test sets balanced by species. This was done to give a sense of the performance of a model that could be trained with the entire dataset, that does not have the species fully balanced. We believe that researchers interested in training models using this dataset for behavior tracking applications would use the entire dataset to fully leverage the variation in the dataset. However, for those interested in training models with balanced species, we provide an annotation file with all the images included, which would allow researchers to create their own training and test sets that meet their specific needs. We have added this justification in the manuscript to guide the other users with different needs. Lines 530-534: “We did not balance our training set for the species as we wanted to utilize the full variation in the dataset and assess models trained with the proportion of species as reflected in the dataset. We provide annotations including the entire dataset to allow others to make create their own training/validation/test sets that suit their needs.”

      And another perhaps major concern that I think should also be addressed somewhere is the fact that this is an image database tested on images while the abstract and manuscript mention the importance of pose estimation for video datasets, yet the current manuscript does not provide any clear test of video datasets nor engage with the practicalities associated with using this image-based database for applications to video datasets. Somewhere this needs to be added to clarify its practical utility.

      We thank the reviewer for this important suggestion. Since we can separate a video into its constituent frames, one can indeed use the provided model or other models trained using this dataset for inference on the frames, thus allowing video tracking applications. We now include a short video clip of a chimpanzee with inferences from the provided model visualized in the supplementary materials.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      Please provide a more thorough description of the annotation procedure (i.e., the instructions given to crowd workers)! See public review for reference on dataset annotation reporting cards.

      We have included the list of instructions for Hive annotators in the supplementary materials.

      An estimate of the crowd worker accuracy and variability would be super valuable!

      While we agree that this is useful, we do not have access to Hive internal data on crowd worker IDs that could allow us to estimate these metrics. Furthermore, we assessed each image manually to ensure good annotation quality.

      In the methods section it is reported that images were discarded because they were either too blurry, small, or highly occluded. Further quantification could be provided. How many images were discarded per species?

      It’s not really clear to us why this is interesting or important. We used a large number of photographers and annotators, some of whom gave a high ratio of great images; some of whom gave a poor ratio. But it’s not clear what those ratios tell us.

      Placing the numerical values at the end of the bars would make the graphs more readable in Figures 4 and 5.

      We thank the reviewer for this suggestion. While we agree that this can help, we do not have space to include the number in a font size that would be readable. Smaller font sizes that are likely to fit may not be readable for all readers. We have included the numerical values in the main text in the results section for those interested and hope that the figures provide a qualitative sense of the results to the readers.

    2. Reviewer #2 (Public Review):

      The authors present the OpenApePose database constituting a collection of over 70000 ape images which will be important for many applications within primatology and the behavioural sciences. The authors have also rigorously tested the utility of this database in comparison to available Pose image databases for monkeys and humans to clearly demonstrate its solid potential. However, the variation in the database with regards to individuals, background, source/setting is not clearly articulated and would be beneficial information for those wishing to make use of this resource in the future. At present, there is also a lack of clarity as to how this image database can be extrapolated to aid video data analyses which would be highly beneficial as well.

      I have two major concerns with regard to the manuscript as it currently stands which I think if addressed would aid the clarity and utility of this database for readers.

      (1) Human annotators are mentioned as doing the 16 landmarks manually for all images but there is no assessment of inter-observer reliability or the such. I think something to this end is currently missing, along with how many annotators there were. This will be essential for others to know who may want to use this database in the future.

      Relevant to this comment, in your description of the database, a table or such could be included, providing the number of images from each source/setting per species and/or number of individuals. Something to give a brief overview of the variation beyond species. (subspecies would also be of benefit for example).

      (2) You mention around line 195 that you used a specific function for splitting up the dataset into training, validation, and test but there is no information given as to whether this was simply random or if an attempt to balance across species, individuals, background/source was made. I would actually think that a balanced approach would be more appropriate/useful here so whether or not this was done, and the reasoning behind that must be justified.

      This is especially relevant given that in one test you report balancing across species (for the sample size subsampling procedure).

      And another perhaps major concern that I think should also be addressed somewhere is the fact that this is an image database tested on images while the abstract and manuscript mention the importance of pose estimation for video datasets, yet the current manuscript does not provide any clear test of video datasets nor engage with the practicalities associated with using this image-based database for applications to video datasets. Somewhere this needs to be added to clarify its practical utility.

    3. Reviewer #1 (Public Review):

      This work provides a new dataset of 71,688 images of different ape species across a variety of environmental and behavioral conditions, along with pose annotations per image. The authors demonstrate the value of their dataset by training pose estimation networks (HRNet-W48) on both their own dataset and other primate datasets (OpenMonkeyPose for monkeys, COCO for humans), ultimately showing that the model trained on their dataset had the best performance (performance measured by PCK and AUC). In addition to their ablation studies where they train pose estimation models with either specific species removed or a certain percentage of the images removed, they provide solid evidence that their large, specialized dataset is uniquely positioned to aid in the task of pose estimation for ape species.

      The diversity and size of the dataset make it particularly useful, as it covers a wide range of ape species and poses, making it particularly suitable for training off-the-shelf pose estimation networks or for contributing to the training of a large foundational pose estimation model. In conjunction with new tools focused on extracting behavioral dynamics from pose, this dataset can be especially useful in understanding the basis of ape behaviors using pose.

      Since the dataset provided is the first large, public dataset of its kind exclusively for ape species, more details should be provided on how the data were annotated, as well as summaries of the dataset statistics. In addition, the authors should provide the full list of hyperparameters for each model that was used for evaluation (e.g., mmpose config files, textual descriptions of augmentation/optimization parameters).

      Overall this work is a terrific contribution to the field and is likely to have a significant impact on both computer vision and animal behavior.

      Strengths: - Open source dataset with excellent annotations on the format, as well as example code provided for working with it. - Properties of the dataset are mostly well described. - Comparison to pose estimation models trained on humans vs monkeys, finding that models trained on human data generalized better to apes than the ones trained on monkeys, in accordance with phylogenetic similarity. This provides evidence for an important consideration in the field: how well can we expect pose estimation models to generalize to new species when using data from closely or distantly related ones? - Sample efficiency experiments reflect an important property of pose estimation systems, which indicates how much data would be necessary to generate similar datasets in other species, as well as how much data may be required for fine-tuning these types of models (also characterized via ablation experiments where some species are left out). - The sample efficiency experiments also reveal important insights about scaling properties of different model architectures, finding that HRNet saturates in performance improvements as a function of dataset size sooner than other architectures like CPMs (even though HRNets still perform better overall).

      Weaknesses: - More details on training hyperparameters used (preferably full config if trained via mmpose). - Should include dataset datasheet, as described in Gebru et al 2021 (arXiv:1803.09010). - Should include crowdsourced annotation datasheet, as described in Diaz et al 2022 (arXiv:2206.08931). Alternatively, the specific instructions that were provided to Hive/annotators would be highly relevant to convey what annotation protocols were employed here. - Should include model cards, as described in Mitchell et al (arXiv:1810.03993). - It would be useful to include more information on the source of the data as they are collected from many different sites and from many different individuals, some of which may introduce structural biases such as lighting conditions due to geography and time of year. - Is there a reason not to use OKS? This incorporates several factors such as landmark visibility, scale, and landmark type-specific annotation variability as in Ronchi & Perona 2017 (arXiv:1707.05388). The latter (variability) could use the human pose values (for landmarks types that are shared), the least variable keypoint class in humans (eyes) as a conservative estimate of accuracy, or leverage a unique aspect of this work (crowdsourced annotations) which affords the ability to estimate these values empirically. - A reporting of the scales present in the dataset would be useful (e.g., histogram of unnormalized bounding boxes) and would align well with existing pose dataset papers such as MS-COCO (arXiv:1405.0312) which reports the distribution of instance sizes and instance density per image.

    4. eLife Assessment

      The OpenApePose dataset presented in this manuscript represents an important contribution to the field of primate behaviour and computer-vision science with methodological applications that are sure to be applicable for a wide variety of taxa. The analysis supporting the utility of this database is solid and compelling but would benefit from some additional clarity, particularly with regards to the annotation of landmarks, model parameters and division of the dataset for training, validation and testing.

    1. eLife Assessment

      This study addresses an important question in aging biology by combining metabolomics, transcriptomics, molecular genetics, and functional analyses to examine how cytosolic acetyl-CoA metabolism influences late-life fitness in replicatively aging yeast. The evidence supporting the roles of AMPK activation, mitochondrial acetyl-CoA utilization, and fatty acid synthesis in preserving fitness during aging is convincing overall, and the engineered A2A strain provides an elegant demonstration that coordinated modulation of distinct acetyl-CoA metabolic branches can increase the proportion of aged cells with a low-senescence phenotype. The study provides significant insight into mechanisms that allow aging cells to maintain fitness without extending replicative lifespan.

    2. Reviewer #1 (Public review):

      This rigorous and creative study uses an elegant combination of metabolomics, transcriptomics, and budding yeast molecular genetics to discover that (i) activating AMPK to maintain mitochondrial respiration fuelled by cytosolic Acetyl CoA and (ii) increasing fatty acid synthesis independent of respiration drive independent pathways that increase the fitness of replicatively-aged budding yeast cells, albeit without increasing their lifespan. The reviewers have achieved their aims and the results support their conclusions. This work provides important insight into molecular mechanisms that allow aging without loss of fitness and will be of interest to scientists in the field of aging and metabolism.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, the authors investigate how cytosolic acetyl-CoA metabolism influences replicative aging in budding yeast. They propose that acetyl-CoA regulates aging through three major pathways: (1) mitochondrial transport to support mitochondrial function, (2) fatty acid synthesis, and (3) global protein acetylation. The data show that AMPK activation promotes mitochondrial import of acetyl-CoA and partially mitigates mitochondrial decline in a subset of aging cells. Furthermore, the engineered A2A strain, which enhances mitochondrial acetyl-CoA utilization while relieving inhibition of fatty acid synthesis, increases the proportion of cells exhibiting a "low senescence" phenotype.

      Overall, this is a thoughtful and potentially impactful study that advances our understanding of metabolic control of aging. Addressing the points below, particularly by refining interpretations and, where feasible, incorporating additional analyses, will further strengthen the manuscript and its conclusions.

      Strengths:

      The study has several notable strengths. It addresses an important question by shifting the focus from lifespan to preservation of late-life fitness, which is highly relevant to aging biology. The work integrates metabolic, genetic, and functional analyses to link cytosolic acetyl-CoA flux with distinct aging outcomes, and the engineering of the A2A strain provides a clear and elegant demonstration of how coordinated pathway modulation can improve cellular fitness.

      Comments on revised version.

      I am fine with the revisions.